The State of Timed Assessment 2026
The annual review of everything PrepClubs Research measured this year. Four reports, nine test formats, and one relationship underneath all of them: what a candidate scores on a timed aptitude test is decided less by what they know than by how many seconds each question is allowed.
Sample 727 re-graded practice attempts from 575 distinct candidates, across 9 formats. 533 of those attempts were on forms allowing under thirty seconds a question.
1Summary of findings
- 57.4% of the final fifth of a speeded paper is never answered against 12.1% where the clock is generous, on tests where a wrong answer costs nothing. Report 2026-01
- Thinking longer about a question is worth 3.1 points once the question is held constant, while the surplus time spent doing it would have covered 7.3 further questions. Report 2026-02
- Papers do not share a difficulty shape. PI Cognitive, CCAT and Wonderlic escalate steeply, Cubiks Logiks and IBEW aptitude run flat, and Watson-Glaser opens with its hardest block. Report 2026-03
- On a speeded paper the raw score is mostly coverage (r=0.72 against r=0.56 for accuracy), and the ordering reverses once the clock stops binding. Report 2026-04
Taken together these describe a single failure mode. A candidate meets a hard item early, spends a minute on it because that feels like the diligent thing to do, gains almost nothing on that item, and pays for it with the last ten questions of a paper that was getting harder anyway and that scores an unanswered item exactly as it scores a wrong one.
2The corpus
Every report in this series is computed from one corpus with one set of filters, so any two of them can be cited side by side. 1,473 attempts are on record. The filters remove attempts that cannot be scored honestly against the form as it now stands, and attempts that show no evidence of having been sat.
| Attempts on record | 1,473 |
| Discarded: re-grade disagreed with the stored score | 556 |
| Discarded: no answers recorded | 26 |
| Discarded: used under half the allowed clock | 164 |
| Analysed | 727 |
| Distinct candidates behind those attempts | 575 |
Candidates who failed to finish are kept, deliberately. Not finishing is the central measurement of the year. Filtering on it would be selecting on the outcome being reported. The full reasoning, including the two thresholds we tested and rejected, is in the method note.
3Finding one: the paper is abandoned at the end
of the final fifth of a speeded aptitude test is left unanswered
Against 12.1% on generous-clock forms. Speeded n=533, generous n=184.
The shortfall is not distributed across the paper. It arrives as a cliff. In the opening fifth 0.2% of items go unanswered; by the fourth fifth it is 31.9%, and in the closing fifth 57.4%. Generous-clock forms, facing the same kinds of reasoning items, end at 12.1%.

Because none of these formats deducts marks for a wrong answer, every one of those blanks is a choice that scores zero with certainty in place of a choice with positive expected value. Read Report 2026-01 in full.
4Finding two: the problem is the budget, not the thinking
The obvious explanation for an abandoned ending is that candidates get stuck. The obvious remedy, repeated in every prep guide, is that dwelling on a question destroys your accuracy on it. We tested that and it is largely untrue.

Compared without controls, accuracy appears to fall by around twenty points from the fastest band to the slowest. Compared inside the same question, where item difficulty is held exactly constant, the gap is 3.1 points, and the faster half wins on only 54% of questions.
The real cost is arithmetic. The median speeded attempt spends 102 seconds beyond a thirty-second-per-item cap, which at a median item time of 14 seconds is 7.3 further questions. 63% of attempts contain at least one question that ate a full minute. Read Report 2026-02 in full.
5Finding three: papers do not share a shape
Candidates are told, as a general rule, that aptitude tests get harder as they go. The rule holds for half the forms measured and fails badly for the rest.

The CCAT form loses 36.4 accuracy points from its opening fifth to its closing fifth and the Wonderlic form 42.8. The IBEW form moves 0.1 points, which is to say it does not move. The Watson-Glaser form runs the other way entirely, opening at 53.1% and closing at 89.7%.
This is why mid-test self-assessment misleads. On a ramping paper the opening block is the easiest block, so feeling comfortable there carries no information. On the Watson-Glaser the reverse applies, and abandoning the paper on the strength of a poor first section is a real and avoidable error. Read Report 2026-03 in full.
6Finding four: on a speeded paper, the score is coverage
A raw score decomposes exactly into how much of the paper was reached multiplied by how well it was answered. Which term dominates is decided by the clock.
| Population | Attempts | Reached | Accuracy on that | r with coverage | r with accuracy |
|---|---|---|---|---|---|
| Speeded forms | 533 | 80% | 72% | 0.72 | 0.56 |
| Generous clock | 184 | 96.1% | 82.1% | 0.59 | 0.77 |
On speeded forms the average attempt answers 72% of what it attempts and reaches 80% of the paper. The missing marks are overwhelmingly on items never seen. Where the clock is generous, coverage rises to 96.1% and stops discriminating between candidates at all, so the score reverts to measuring what people assume it measures. Read Report 2026-04 in full.
7The relationship the year reduces to
Plot every form by the seconds it allows per question against the share of its items that never get an answer, and the four reports collapse into one picture.

The exception is the most instructive point on the chart. McQuaig allows 18 seconds a question, exactly what the CCAT allows, and leaves 0.2% of its items blank against the CCATβs 20.6%. Seconds per question is not the variable. Seconds per unit of work is, and a candidate cannot read that off the instructions: they have to have sat the format before.
8Every form at a glance

| Form | Items | Seconds each | Attempts | Left blank | Gradient | Reached | Median score |
|---|---|---|---|---|---|---|---|
| PI Cognitive | 50 | 14 | 215 | 18.7% | +31.6 pts | 81.3% | 56% |
| CCAT | 50 | 18 | 189 | 20.6% | +36.4 pts | 79.4% | 54% |
| Cubiks Logiks | 50 | 14 | 76 | 21.9% | +10 pts | 78.1% | 68% |
| IBEW aptitude | 69 | 84 | 61 | 7.6% | -0.1 pts | 92.4% | 87% |
| Wonderlic | 50 | 14 | 53 | 20% | +42.8 pts | 80% | 52% |
| Watson-Glaser | 40 | 45 | 41 | 0.4% | -36.6 pts | 99.6% | 80% |
| EIAT | 90 | 60 | 29 | 2.7% | β | β | β |
| McQuaig | 50 | 18 | 25 | 0.2% | β | β | β |
| Ramsay Mechanical | 36 | 33 | 17 | 4.9% | β | β | β |
A dash means the form cleared the fifteen-attempt floor for the abandonment figures but not the thirty-attempt floor for the gradient and score figures.
9What we withdrew, and what we will not publish
Withdrawn, 21 September 2026. We published, and retracted within a day, the claim that time pressure degrades the answers a candidate does give. Splitting speeded attempts by how much of the clock they actually used killed it: candidates who finished with time to spare declined 27.6 accuracy points against 31.3 for the time-pressured. Almost the whole effect was present in people under no pressure at all. What was left, once that was taken seriously, became Report 2026-03, which is the better finding.
Measured and not published. Attempts that were re-sat show a gain of roughly twelve points. We are not publishing it as a finding about practice, because 61 of the 80 repeat pairs re-sat the same form, so most of the gain is item memory rather than improvement. It is the one number in this corpus that would flatter the company that collected it, which is precisely why it stays where it is until a clean comparison exists.
Not published at all. The per-question records. Joined to an attempt, timing data at this sample size is re-identifiable. Every report publishes its aggregated figures as CSV and none of them publishes the underlying rows.
10Standards this department holds itself to
- Every discard is published. The ledger in section 2 appears in identical form on every report. More than half the attempts on record are excluded, and the reason is stated rather than buried.
- No filtering on the outcome. The filters are fixed before a finding is looked at, and a filter that would select on the thing being measured is rejected even when it produces a cleaner number.
- A confound gets a control, not a caveat. Both of this yearβs main findings survive a direct control: within-item for pacing, full-coverage-only for the ramp. A caveat is what we write when a control is not possible, and then the claim is weakened to match.
- Withdrawals stay visible. A retracted finding is documented with its reason on the report that replaced it, not deleted.
- Figures are files. Every chart is generated by the same script that writes the numbers, so a figure cannot drift from its table, and both the image and a vector copy are downloadable.
11Glossary
- Speeded form
- A test that allows under thirty seconds per question, so the binding constraint is the clock rather than the difficulty. On this platform: the CCAT, Wonderlic, PI Cognitive and Cubiks Logiks formats.
- Generous clock
- A test allowing thirty-three seconds or more per question, where nearly every candidate reaches nearly every item. On this platform: IBEW, Watson-Glaser, EIAT and Ramsay Mechanical formats.
- Coverage
- The share of items on the paper that received any answer at all, whether right or wrong.
- Blank, or unanswered
- An item that received no answer. Every format measured here scores a blank identically to a wrong answer, which is to say zero, with no deduction for the error.
- Gradient
- Accuracy in the opening fifth of the paper minus accuracy in the closing fifth, in percentage points. Positive means the paper gets harder as it goes.
- Full-coverage attempt
- An attempt that answered every single item. Used as the control population for the gradient, because no candidate is missing from any point on the curve.
- Within-item comparison
- A comparison drawn inside a single question, splitting its sittings at that question's own median time, so item difficulty is held exactly constant.
- Re-grade filter
- Every attempt is scored again against the form as it now exists on disk. Disagreements are discarded, because banks were edited after launch and an old attempt cannot be scored against a paper it was not sat on.
12The forms this report is computed from
Every form here is a PrepClubs practice paper written to that test's published format and item mix. The pages below cover the format, and each carries a free practice paper.
- PI Cognitive Assessment
14 seconds a question over 50 items. 18.7% of items left blank across 215 analysed attempts.
- CCAT
18 seconds a question over 50 items. 20.6% of items left blank across 189 analysed attempts.
- Cubiks Logiks
14 seconds a question over 50 items. 21.9% of items left blank across 76 analysed attempts.
- IBEW aptitude test
84 seconds a question over 69 items. 7.6% of items left blank across 61 analysed attempts.
- Wonderlic
14 seconds a question over 50 items. 20% of items left blank across 53 analysed attempts.
- Watson-Glaser
45 seconds a question over 40 items. 0.4% of items left blank across 41 analysed attempts.
- EIAT
60 seconds a question over 90 items. 2.7% of items left blank across 29 analysed attempts.
- McQuaig Mental Agility
18 seconds a question over 50 items. 0.2% of items left blank across 25 analysed attempts.
- Ramsay Mechanical
33 seconds a question over 36 items. 4.9% of items left blank across 17 analysed attempts.
13Common questions
- What is the single most useful thing to know before a timed aptitude test?
- That an unanswered question is the only response on the paper guaranteed to score nothing. Across 727 analysed attempts, none of the nine formats measured deducts marks for a wrong answer, yet 57.4% of the final fifth of a speeded paper is left blank. Stopping with thirty seconds left and filling every blank is worth roughly two marks on a fifty-question paper, for no preparation at all.
- Which aptitude tests are hardest to finish?
- The ones that allow least time per unit of work, which is not the same as least time per question. The CCAT, Wonderlic, PI Cognitive and Cubiks Logiks formats all allow under twenty seconds an item and run between 18.7% and 21.9% blank. The McQuaig format allows the same eighteen seconds as the CCAT and runs at 0.2%, because its items are quick association rather than multi-step reasoning.
- Is it better to answer fewer questions carefully or more questions quickly?
- More questions. On speeded forms the raw score correlates with how much of the paper was reached at r=0.72 and with accuracy on what was reached at r=0.56. Where the clock is generous that ordering reverses, to r=0.77 for accuracy against r=0.59 for coverage. The right strategy is a property of the clock, not a matter of temperament.
- Do these findings come from real employer assessments?
- No. They come from 727 practice attempts on PrepClubs forms, written to each test's published format and sat voluntarily. They are not official scores from any publisher, they are not scaled or normed, and the stakes are not those of a job application. Every report states this where a number is given.
- Can I reuse these figures?
- Yes, with attribution to PrepClubs Research and a link to the report the figure comes from. Every report carries a suggested citation, a CSV of its figures and a vector copy of each chart. The underlying per-question records are not published, because joined to an attempt they are re-identifiable at this sample size.
14Being measured next
- Where the clock actually goes. The distribution of time across the paper rather than across items, and whether candidates who front-load their clock can be identified before they run out.
- Item-level difficulty calibration. Ranking individual questions by observed difficulty rather than by position, which would separate a genuine ramp from an ordering artefact.
- A clean repeat comparison. A second sitting on a different form of the same test, which is the only design that can separate practice from item memory. It requires data we do not yet have.
If you want a cut we have not published, or a figure checked against the source, ask and we will run it.
Suggested citation
βOn speeded pre-employment aptitude tests, 57.4% of the final fifth of the paper is never answered, against 12.1% where the clock is generous. The lost marks are not lost to difficulty: holding the question constant, slower answers are only 3.1 points worse than faster ones, while the median attempt spends enough surplus time on long items to cover 7.3 further questions. PrepClubs Research, PCR-2026.β
Download the figures as CSVMethod and standardsAsk us to check a figure
Reproduce any figure or table with attribution to PrepClubs Research and a link to this page. The underlying per-question records are not published: joined to an attempt they are re-identifiable at this sample size.
Other reports in this series
- 2026-01: How much of a timed aptitude test is never answered
Abandonment by position on speeded and generous-clock forms.
- 2026-02: What an extra thirty seconds on a question actually buys
The return on deliberation, measured inside the same question.
- 2026-03: Which aptitude tests are built as a difficulty ramp
Accuracy by position, with selection controlled out.
- 2026-04: What candidates actually score on aptitude practice tests
Percentile bands per form, and what a raw score is made of.