We have corrected an error in the regression coefficient that converts AvrFreqRank (average frequency rank) into CEFR units.
What was wrong. The correct coefficient derived from the calibration data is 0.010821 (intercept −2.802030), but the released version used 0.004493 (intercept −0.608127). As a result, AvrFreqRank was reported lower than it should have been, and it advanced by only 0.41 per CEFR level where the other seven measures advance by about 1.0.
Effect on the predicted level. CVLA3 predicts a level by discarding the highest and lowest of the eight measures and averaging the remaining six. The under-reported AvrFreqRank was the discarded minimum for about 40% of the texts in our validation data, so for those it had no effect on the predicted level at all. Overall accuracy is almost unchanged: on our development validation data, accuracy went from 65.7% to 63.0%, while agreement within one level improved from 99.1% to 100%. Individual texts can still change, however: the CEFR-J level changed for about a quarter of them.
When this matters to you. If you read the individual measures or their chart, AvrFreqRank will now appear higher than before. Many cases where vocabulary frequency looked lower than the other measures were an artefact of this error. Please take this into account when comparing per-measure results with earlier output. If you have results for a particular text from an earlier version, please re-run it rather than assuming the level is unchanged.
Recalibration of the Listening monologue mode. The listening adjustment had been derived from listening texts scored with the old coefficient, so it had to be re-derived. Version 3.1 added a fixed amount to the regression score (+0.6 when the score was 1 or below, +1.1 otherwise). Version 3.2 instead maps the listening score onto the reading scale with a linear transformation:
Adjusted score = 1.0897 × Regression score + 0.2876
This was fitted on level-graded listening and reading corpora (A1, A2, B1 and B2; 9,000–43,000 words per file) by regressing the score of each listening file onto the score of the reading file at the same CEFR level (r² = .967). All four listening files are assigned their correct level under the new equation. Listening results therefore differ from version 3.1, generally coming out lower, because the earlier adjustment was partly compensating for the under-reported AvrFreqRank. Reading results are not affected by this change.
We apologise for the inconvenience, and we welcome questions and reports via this form.
Uchida, S., & Negishi, M. (2018) Assigning CEFR-J levels to English texts based on textual features. In Y. Tono and H. Isahara (eds.) Proceedings of the 4th Asia Pacific Corpus Linguistics Conference (APCLC 2018), pp. 463-467. [PDF]
内田諭・根岸雅史(2021)「英語読解教材のCEFRレベルの推定 : CVLAの妥当性評価」Journal of Corpus-based Lexicology Studies, 3, pp.1-14. [Link]
Uchida S., & Negishi, M. (2025) Estimating the CEFR-J level of English reading passages: Development and accuracy of CVLA3『英語コーパス研究』(English Corpus Studies) 32. pp.165-174. https://doi.org/10.69193/ecs.32.0_165 [Link]
Aug 21, 2026: CVLA 3.2 released. Corrected the AvrFreqRank regression coefficient (0.004493 → 0.010821, intercept −0.608127 → −2.802030).
Aug 21, 2026: Recalibrated the Listening monologue adjustment to 1.0897 × score + 0.2876 (previously +0.6 / +1.1).
Aug 21, 2026: CVLA3 Desktop 3.2 released for Windows and macOS, with the same correction, the Reading/Listening monologue mode selector, and results aligned with the online version (identical values for all eight measures). Version 3.0 (2025) remains downloadable for reproducing earlier analyses.
Mar 25, 2026: CVLA 3.1 released. Added Reading/Listening monologue selector and listening-specific CEFR conversion adjustment.
May 7, 2025: The UI of this page has been updated.
May 6, 2025: The desktop version of CVLA3 (beta) released.
Feburary 9, 2025: CVLA 3.0 released.
October 23, 2024: CVLA 3.0 beta released.