PrepClubs ResearchReport PCR-2026Edition 1Published

The State of Timed Assessment 2026

The annual review of everything PrepClubs Research measured this year. Four reports, nine test formats, and one relationship underneath all of them: what a candidate scores on a timed aptitude test is decided less by what they know than by how many seconds each question is allowed.

Sample 727 re-graded practice attempts from 575 distinct candidates, across 9 formats. 533 of those attempts were on forms allowing under thirty seconds a question.

1Summary of findings

  1. 57.4% of the final fifth of a speeded paper is never answered against 12.1% where the clock is generous, on tests where a wrong answer costs nothing. Report 2026-01
  2. Thinking longer about a question is worth 3.1 points once the question is held constant, while the surplus time spent doing it would have covered 7.3 further questions. Report 2026-02
  3. Papers do not share a difficulty shape. PI Cognitive, CCAT and Wonderlic escalate steeply, Cubiks Logiks and IBEW aptitude run flat, and Watson-Glaser opens with its hardest block. Report 2026-03
  4. On a speeded paper the raw score is mostly coverage (r=0.72 against r=0.56 for accuracy), and the ordering reverses once the clock stops binding. Report 2026-04

Taken together these describe a single failure mode. A candidate meets a hard item early, spends a minute on it because that feels like the diligent thing to do, gains almost nothing on that item, and pays for it with the last ten questions of a paper that was getting harder anyway and that scores an unanswered item exactly as it scores a wrong one.

2The corpus

Every report in this series is computed from one corpus with one set of filters, so any two of them can be cited side by side. 1,473 attempts are on record. The filters remove attempts that cannot be scored honestly against the form as it now stands, and attempts that show no evidence of having been sat.

Attempts on record1,473
Discarded: re-grade disagreed with the stored score556
Discarded: no answers recorded26
Discarded: used under half the allowed clock164
Analysed727
Distinct candidates behind those attempts575

Candidates who failed to finish are kept, deliberately. Not finishing is the central measurement of the year. Filtering on it would be selecting on the outcome being reported. The full reasoning, including the two thresholds we tested and rejected, is in the method note.

3Finding one: the paper is abandoned at the end

57.4%

of the final fifth of a speeded aptitude test is left unanswered

Against 12.1% on generous-clock forms. Speeded n=533, generous n=184.

The shortfall is not distributed across the paper. It arrives as a cliff. In the opening fifth 0.2% of items go unanswered; by the fourth fifth it is 31.9%, and in the closing fifth 57.4%. Generous-clock forms, facing the same kinds of reasoning items, end at 12.1%.

Unanswered questions on timed aptitude tests, by position in the paper. Speeded forms rise from 0.2% left blank in the opening fifth to 57.4% in the closing fifth; generous-clock forms rise only to 12.1%.
Figure 1. Reproduced from Report 2026-01. The gap between the two lines is the clock. Vector copy

Because none of these formats deducts marks for a wrong answer, every one of those blanks is a choice that scores zero with certainty in place of a choice with positive expected value. Read Report 2026-01 in full.

4Finding two: the problem is the budget, not the thinking

The obvious explanation for an abandoned ending is that candidates get stuck. The obvious remedy, repeated in every prep guide, is that dwelling on a question destroys your accuracy on it. We tested that and it is largely untrue.

Does spending longer on an aptitude test question help? Inside the same question, the faster half of sittings scores 72.5% and the slower half 69.4%, a gap of 3.1 points across 190 questions.
Figure 2. Reproduced from Report 2026-02. Across 190 questions sat at least 20 times, split at each question's own median time. Vector copy

Compared without controls, accuracy appears to fall by around twenty points from the fastest band to the slowest. Compared inside the same question, where item difficulty is held exactly constant, the gap is 3.1 points, and the faster half wins on only 54% of questions.

The real cost is arithmetic. The median speeded attempt spends 102 seconds beyond a thirty-second-per-item cap, which at a median item time of 14 seconds is 7.3 further questions. 63% of attempts contain at least one question that ate a full minute. Read Report 2026-02 in full.

5Finding three: papers do not share a shape

Candidates are told, as a general rule, that aptitude tests get harder as they go. The rule holds for half the forms measured and fails badly for the rest.

Which aptitude tests get harder as they go. Accuracy points lost from the opening fifth to the closing fifth, from 42.8 on the Wonderlic form and 36.4 on the CCAT down to -36.6 on the Watson-Glaser, which opens with its hardest block.
Figure 3. Reproduced from Report 2026-03. Positive means the paper gets harder as it goes. Vector copy

The CCAT form loses 36.4 accuracy points from its opening fifth to its closing fifth and the Wonderlic form 42.8. The IBEW form moves 0.1 points, which is to say it does not move. The Watson-Glaser form runs the other way entirely, opening at 53.1% and closing at 89.7%.

This is why mid-test self-assessment misleads. On a ramping paper the opening block is the easiest block, so feeling comfortable there carries no information. On the Watson-Glaser the reverse applies, and abandoning the paper on the strength of a poor first section is a real and avoidable error. Read Report 2026-03 in full.

6Finding four: on a speeded paper, the score is coverage

A raw score decomposes exactly into how much of the paper was reached multiplied by how well it was answered. Which term dominates is decided by the clock.

Table 1. What the raw score follows
PopulationAttemptsReachedAccuracy on thatr with coverager with accuracy
Speeded forms53380%72%0.720.56
Generous clock18496.1%82.1%0.590.77

On speeded forms the average attempt answers 72% of what it attempts and reaches 80% of the paper. The missing marks are overwhelmingly on items never seen. Where the clock is generous, coverage rises to 96.1% and stops discriminating between candidates at all, so the score reverts to measuring what people assume it measures. Read Report 2026-04 in full.

7The relationship the year reduces to

Plot every form by the seconds it allows per question against the share of its items that never get an answer, and the four reports collapse into one picture.

Seconds allowed per question on each aptitude test against the share of the paper left unanswered, one point per form. The four forms under twenty seconds a question cluster between 18.7% and 21.9% blank, while forms allowing thirty seconds or more sit under 7.6%.
Figure 4. One point per form. The dashed line marks thirty seconds a question, below which the clock binds. Vector copy

The exception is the most instructive point on the chart. McQuaig allows 18 seconds a question, exactly what the CCAT allows, and leaves 0.2% of its items blank against the CCAT’s 20.6%. Seconds per question is not the variable. Seconds per unit of work is, and a candidate cannot read that off the instructions: they have to have sat the format before.

8Every form at a glance

How much of each aptitude test paper candidates reach, against the median score they get. CCAT, Wonderlic, PI Cognitive and Cubiks Logiks all show a large gap between the two bars; the IBEW aptitude and Watson-Glaser papers show almost none. Where the bars diverge, the clock rather than the difficulty is limiting the score.
Figure 5. A paper you run out of and a paper you find hard look nothing alike. Vector copy
Table 2. Summary of every form with at least fifteen analysed attempts
FormItemsSeconds eachAttemptsLeft blankGradientReachedMedian score
PI Cognitive501421518.7%+31.6 pts81.3%56%
CCAT501818920.6%+36.4 pts79.4%54%
Cubiks Logiks50147621.9%+10 pts78.1%68%
IBEW aptitude6984617.6%-0.1 pts92.4%87%
Wonderlic50145320%+42.8 pts80%52%
Watson-Glaser4045410.4%-36.6 pts99.6%80%
EIAT9060292.7%β€”β€”β€”
McQuaig5018250.2%β€”β€”β€”
Ramsay Mechanical3633174.9%β€”β€”β€”

A dash means the form cleared the fifteen-attempt floor for the abandonment figures but not the thirty-attempt floor for the gradient and score figures.

9What we withdrew, and what we will not publish

Withdrawn, 21 September 2026. We published, and retracted within a day, the claim that time pressure degrades the answers a candidate does give. Splitting speeded attempts by how much of the clock they actually used killed it: candidates who finished with time to spare declined 27.6 accuracy points against 31.3 for the time-pressured. Almost the whole effect was present in people under no pressure at all. What was left, once that was taken seriously, became Report 2026-03, which is the better finding.

Measured and not published. Attempts that were re-sat show a gain of roughly twelve points. We are not publishing it as a finding about practice, because 61 of the 80 repeat pairs re-sat the same form, so most of the gain is item memory rather than improvement. It is the one number in this corpus that would flatter the company that collected it, which is precisely why it stays where it is until a clean comparison exists.

Not published at all. The per-question records. Joined to an attempt, timing data at this sample size is re-identifiable. Every report publishes its aggregated figures as CSV and none of them publishes the underlying rows.

10Standards this department holds itself to

  • Every discard is published. The ledger in section 2 appears in identical form on every report. More than half the attempts on record are excluded, and the reason is stated rather than buried.
  • No filtering on the outcome. The filters are fixed before a finding is looked at, and a filter that would select on the thing being measured is rejected even when it produces a cleaner number.
  • A confound gets a control, not a caveat. Both of this year’s main findings survive a direct control: within-item for pacing, full-coverage-only for the ramp. A caveat is what we write when a control is not possible, and then the claim is weakened to match.
  • Withdrawals stay visible. A retracted finding is documented with its reason on the report that replaced it, not deleted.
  • Figures are files. Every chart is generated by the same script that writes the numbers, so a figure cannot drift from its table, and both the image and a vector copy are downloadable.

11Glossary

Speeded form
A test that allows under thirty seconds per question, so the binding constraint is the clock rather than the difficulty. On this platform: the CCAT, Wonderlic, PI Cognitive and Cubiks Logiks formats.
Generous clock
A test allowing thirty-three seconds or more per question, where nearly every candidate reaches nearly every item. On this platform: IBEW, Watson-Glaser, EIAT and Ramsay Mechanical formats.
Coverage
The share of items on the paper that received any answer at all, whether right or wrong.
Blank, or unanswered
An item that received no answer. Every format measured here scores a blank identically to a wrong answer, which is to say zero, with no deduction for the error.
Gradient
Accuracy in the opening fifth of the paper minus accuracy in the closing fifth, in percentage points. Positive means the paper gets harder as it goes.
Full-coverage attempt
An attempt that answered every single item. Used as the control population for the gradient, because no candidate is missing from any point on the curve.
Within-item comparison
A comparison drawn inside a single question, splitting its sittings at that question's own median time, so item difficulty is held exactly constant.
Re-grade filter
Every attempt is scored again against the form as it now exists on disk. Disagreements are discarded, because banks were edited after launch and an old attempt cannot be scored against a paper it was not sat on.

12The forms this report is computed from

Every form here is a PrepClubs practice paper written to that test's published format and item mix. The pages below cover the format, and each carries a free practice paper.

13Common questions

What is the single most useful thing to know before a timed aptitude test?
That an unanswered question is the only response on the paper guaranteed to score nothing. Across 727 analysed attempts, none of the nine formats measured deducts marks for a wrong answer, yet 57.4% of the final fifth of a speeded paper is left blank. Stopping with thirty seconds left and filling every blank is worth roughly two marks on a fifty-question paper, for no preparation at all.
Which aptitude tests are hardest to finish?
The ones that allow least time per unit of work, which is not the same as least time per question. The CCAT, Wonderlic, PI Cognitive and Cubiks Logiks formats all allow under twenty seconds an item and run between 18.7% and 21.9% blank. The McQuaig format allows the same eighteen seconds as the CCAT and runs at 0.2%, because its items are quick association rather than multi-step reasoning.
Is it better to answer fewer questions carefully or more questions quickly?
More questions. On speeded forms the raw score correlates with how much of the paper was reached at r=0.72 and with accuracy on what was reached at r=0.56. Where the clock is generous that ordering reverses, to r=0.77 for accuracy against r=0.59 for coverage. The right strategy is a property of the clock, not a matter of temperament.
Do these findings come from real employer assessments?
No. They come from 727 practice attempts on PrepClubs forms, written to each test's published format and sat voluntarily. They are not official scores from any publisher, they are not scaled or normed, and the stakes are not those of a job application. Every report states this where a number is given.
Can I reuse these figures?
Yes, with attribution to PrepClubs Research and a link to the report the figure comes from. Every report carries a suggested citation, a CSV of its figures and a vector copy of each chart. The underlying per-question records are not published, because joined to an attempt they are re-identifiable at this sample size.

14Being measured next

  • Where the clock actually goes. The distribution of time across the paper rather than across items, and whether candidates who front-load their clock can be identified before they run out.
  • Item-level difficulty calibration. Ranking individual questions by observed difficulty rather than by position, which would separate a genuine ramp from an ordering artefact.
  • A clean repeat comparison. A second sitting on a different form of the same test, which is the only design that can separate practice from item memory. It requires data we do not yet have.

If you want a cut we have not published, or a figure checked against the source, ask and we will run it.

Suggested citation

β€œOn speeded pre-employment aptitude tests, 57.4% of the final fifth of the paper is never answered, against 12.1% where the clock is generous. The lost marks are not lost to difficulty: holding the question constant, slower answers are only 3.1 points worse than faster ones, while the median attempt spends enough surplus time on long items to cover 7.3 further questions. PrepClubs Research, PCR-2026.”

Download the figures as CSVMethod and standardsAsk us to check a figure

Reproduce any figure or table with attribution to PrepClubs Research and a link to this page. The underlying per-question records are not published: joined to an attempt they are re-identifiable at this sample size.

Other reports in this series