Reference
What we assess against
This page documents the framework our reviewers apply. It is written as a working reference, not as marketing — teachers are welcome to use it even if they never send us anything.
1. CEFR level calibration
The Common European Framework of Reference (CEFR) describes language ability on a six-level scale, A1–C2. When a test declares a level, every part of it makes a claim: the texts, the target items, the distractors, and the instructions all have to sit at — or deliberately below — that level. Our calibration check tests that claim.
In practice we check four layers, in this order:
- Target vocabulary. The words an item actually tests are referenced against public CEFR vocabulary inventories. A B1 gap that can only be filled by notwithstanding is a C1 item wearing a B1 label.
- Carrier language. The text around the items. Carrier language may sit below the declared level; it must not sit above it, or the item measures reading of the carrier rather than the target point.
- Grammar inventory. Structures tested are checked against where they are commonly introduced. Boundary cases (say, second conditional at A2/B1) are not failures; they are flagged in the calibration record so the test owner can decide.
- Task demand. A text can be B1 while the task is C1 — inference across paragraphs, recognising attitude, or synthesising two viewpoints raise the level even when every word is simple.
The output of this check is not a pass/fail stamp but a level evidence table: for each section, what the vocabulary, grammar, and task demand indicate, and where that disagrees with the declared level.
2. Task-format conformance
Many client tests announce a relationship to a public exam format — “in the style of a B2 First Reading and Use of English, Part 2”, “IELTS-style Task 1 report”. We treat such a statement as a specification and verify against the exam board’s own published format description: number of items, text length, task type, and instruction wording. The reference is descriptive; we are independent of all exam boards, and conformance with a format is not endorsement by its owner.
Typical conformance findings:
- An “open cloze” with nine gaps where the referenced part specifies eight — students practising timing learn the wrong rhythm.
- A gapped text of 210 words where the referenced format runs 140–190 — every timing assumption built on it is off.
- Multiple-choice items with three options where the format uses four, silently raising the guessing baseline from 25% to 33%.
- Instruction wording that differs from the format’s standard rubric — a small thing, until a student meets the real rubric for the first time in the exam room.
Where a test declares no external format, this check is skipped and we assess internal consistency instead: parallel sections behave the same way throughout the paper.
3. Item-writing principles
Our reviews apply the same principles professional item writers work to. The ones that catch the most findings:
- One defensible key. Exactly one option may be correct, under any reasonable reading of the stem. The test writer cannot check this alone: knowing the intended answer makes the second defensible option invisible. This is category 2 in our reports and our most frequent high-severity finding.
- Distractors must work for their living. Each wrong option should be chosen by somebody who has not mastered the point. Options nobody picks are dead weight that shortens the real test.
- The text must be necessary. If a comprehension item can be answered from world knowledge or by matching surface words, it does not measure comprehension.
- Independence of items. Getting item 4 wrong must not make item 5 unanswerable.
- No trick items. Difficulty should come from the language point, not from ambiguity, obscure trivia, or deliberately misleading stems.
4. Severity definitions
Every finding carries one of three severities. The definitions are fixed, so severities are comparable across reports and reviewers:
| Severity | Definition | Example |
|---|---|---|
| High | The item cannot be scored fairly as it stands: wrong key, two defensible answers, or unanswerable stem. | Key marks B; C is also correct. |
| Medium | The item scores, but measures the wrong thing or the wrong level: guessable distractors, off-level vocabulary, format drift. | Reading item answerable without the text. |
| Low | Does not affect scoring: typos, punctuation, US/UK inconsistency, rubric wording. | organize in an otherwise UK-convention paper. |
5. Language conventions (US/UK)
We do not treat either convention as correct; we treat inconsistency as the error. A paper is checked against the convention it predominantly uses — spelling (colour / color), punctuation around quotation marks, date formats, and vocabulary pairs (autumn / fall, timetable / schedule). Where a test prepares for a specific exam, the convention that exam’s materials use is the default, and we say so in the finding rather than silently imposing it.
6. Independence statement
EnglishCorrection is not affiliated with, endorsed by, or accredited by any examination board, university, or publisher. References on this page and in our reports to public exam formats and to the CEFR are descriptive. Our certificate attests that an independent review was carried out on a specific version of a specific document — nothing more, and we consider that exactly enough. Reviews are carried out by named reviewers, and our fee never depends on the review’s outcome.
Want this framework applied to your test?
Send the material and the declared level. The findings report references the sections above, so you can check our reasoning.
