, 5 minutes to read
Every school wants a placement test that is short, because candidates abandon long ones, and exact, because wrong placements cost students. The two pull against each other. This guide shows where the balance lies, with figures from our own question bank.
Why longer is not simply better
Each extra question adds a little information and takes a little more of the candidate’s patience. The information added shrinks with every question: the tenth tells you far more than the fortieth. Patience runs out the other way. Somewhere the curves cross, and beyond that point a longer test is a worse one, because tired candidates answer carelessly or leave.
What the figures say
We simulated 3,000 candidates, with the same spread of ability as our real ones, taking an adaptive test on our English question bank at different lengths. For each length we counted how often the result was the candidate’s true level out of eight (A1 to C2 with B1+ and B2+), and how often it was that level or the one next to it.
| Scored questions | Exactly the true level | True level or the next |
|---|---|---|
| 15 | 55% | 96% |
| 20 | 57% | 98% |
| 25 | 61% | 99% |
| 30 | 64% | 99% |
| 45 | 72% | 100% |
Three things stand out.
- Gross errors disappear early. With fifteen well-chosen questions, 96 candidates in 100 are on their level or the one beside it. By twenty-five it is 99.
- Exactness comes slowly. Tripling the test from fifteen to forty-five questions lifts the share placed exactly from 55% to 72%. That is a real gain, bought with thirty more questions.
- No length reaches 100% exact. The reason is not the test but the scale, as the next section explains.
These are results of a simulation: it assumes that the measured difficulty of each question is right. Real candidates add noise that a simulation does not have, such as slips, distraction and good guesses.
The boundary problem
Levels are bands on a continuous scale, and candidates are spread along it. A good share of them always stand close to a line between two levels. For those candidates “B1 or B1+” is not a failure of measurement; it is the truth. A hundred more questions would only confirm that they are on the line.
This is why a level on its own is a poor result, and why a score with a margin is a better one. It tells you which candidates are safely inside a level and which deserve two minutes of conversation.
So how long?
For a placement decision, as opposed to an exam, the evidence points to a simple rule:
- Around twenty to thirty adaptive questions are enough to rule out gross errors for practically everyone.
- Let the test stop early for candidates who are clearly inside a level, and run longer for those near a boundary.
- Spend the time you saved on people: a short conversation with the borderline cases adds information that no number of extra grammar questions can.
NGPT asks between 18 and 30 scored questions and stops when it is sure enough. How it decides is described under How it works.
A fixed test needs to be much longer to reach the same precision, because most of its questions are far from the candidate’s level. See Adaptive testing explained.