PrepClubs ResearchReport 2026-03Edition 1Published

Which aptitude tests are built as a difficulty ramp

Candidates are routinely told that aptitude tests get harder as they go. Across the forms measured here that is true of half of them, false of two, and backwards on one. The shape of the paper is a property of the paper, and a candidate who does not know which shape they are sitting will misread their own progress.

Sample 727 re-graded practice attempts from 575 distinct candidates. 6 forms cleared the 30-attempt floor for inclusion.

1Finding

36.4 vs -36.6 points

separates the steepest ramp from the steepest reverse ramp

Accuracy lost from the opening fifth of the paper to the closing fifth. A positive figure means the paper gets harder; a negative one means it gets easier.

Of the 6 forms with enough attempts to measure, 3 escalate steeply (PI Cognitive, CCAT and Wonderlic), 2 run essentially flat (Cubiks Logiks and IBEW aptitude), and one runs backwards (Watson-Glaser), with its hardest block at the very start.

This matters because candidates calibrate mid-test. A paper that felt manageable for ten minutes is read as evidence of a good performance, and on a steep ramp that reading is wrong. The opening of a CCAT form is answered at 81.7% and the close at 45.3%.

2Three shapes, not one

Do aptitude tests get harder as they go? Accuracy points lost between the opening fifth and the closing fifth of each paper: Wonderlic +42.8, CCAT +36.4, PI Cognitive +31.6, Cubiks Logiks +10, IBEW aptitude -0.1, Watson-Glaser -36.6. A positive figure means the paper escalates; a negative one means it opens with its hardest block.
Figure 1. Positive means the paper gets harder as it goes. Negative means it gets easier. Computed on attempts that answered every item wherever the count allows. Vector copy
Table 1. Accuracy gradient, opening fifth to closing fifth
FormClockAttemptsGradientBasis
Wonderlic14s a question53+42.8 ptsanswered items
CCAT18s a question189+36.4 ptsfull coverage
PI Cognitive14s a question215+31.6 ptsfull coverage
Cubiks Logiks14s a question76+10 ptsfull coverage
IBEW aptitude84s a question61-0.1 ptsfull coverage
Watson-Glaser45s a question41-36.6 ptsfull coverage

"Full coverage" means the gradient was computed only on attempts that answered every item. Where fewer than 20 such attempts exist, the answered-item basis is used instead and is noted as such.

3The curves themselves

A gradient reduces a curve to one number, which hides whether a paper climbs steadily or steps once. Figure 2 shows the curves for every form with enough full-coverage attempts to draw one honestly.

Accuracy by position in the paper for CCAT, PI Cognitive, Cubiks Logiks, IBEW aptitude, Watson-Glaser, on attempts that answered every question. The CCAT and PI Cognitive curves fall steeply, the IBEW aptitude curve is flat, and the Watson-Glaser curve rises sharply after its opening block.
Figure 2. A flat line is a paper of even difficulty. A falling line is a ramp. The Watson-Glaser line shows a paper that opens with its hardest work. Vector copy
Table 2. Accuracy by fifth of the paper
FormBasis1st2nd3rd4th5th
Wonderlicanswered items82.5%75.9%63.3%58.9%39.7%
CCATfull coverage81.7%72.8%73%56.6%45.3%
PI Cognitivefull coverage87.7%65%66.4%57.3%56.1%
Cubiks Logiksfull coverage86.8%77.7%81.4%82.7%76.8%
IBEW aptitudefull coverage86.4%86.9%84.7%87.5%86.5%
Watson-Glaserfull coverage53.1%85.3%81.9%90.9%89.7%

PI Cognitive steps once and then holds. Its opening fifth runs at 87.7% and its second at 65%, a fall of 22.7 points inside the first two blocks, after which it drifts. The CCAT, by contrast, declines throughout. Two papers with the same overall gradient can demand entirely different pacing.

4Controlling for who reaches the end

There is an obvious objection to measuring accuracy by position on a speeded form. The later items are answered only by candidates fast enough to reach them, who are on average the stronger candidates, so late accuracy is measured on a self-selected group.

That objection cuts the wrong way. Selection of this kind should push late accuracy up, not down. Finding a steep decline despite it means the true gradient is at least as steep as reported.

The control is direct rather than statistical: recompute on attempts that answered every single item, where no candidate is missing from any point on the curve. Both columns are published in Table 3 and they agree.

Table 3. Gradient before and after removing selection
FormAll answered itemsFull-coverage attempts onlyn (full coverage)
PI Cognitive31.9 pts31.6 pts70
CCAT35.1 pts36.4 pts47
Cubiks Logiks9.4 pts10.0 pts22
IBEW aptitude0.1 pts-0.1 pts49
Wonderlic42.8 ptstoo few
Watson-Glaser-37.3 pts-36.6 pts40

The Wonderlic form has fewer than 20 full-coverage attempts, so its gradient is reported on the answered-item basis and should be read as a lower bound.

5What it changes for a candidate

  • Do not calibrate on the opening. On a ramping form the first ten questions are the easiest ten questions. Feeling comfortable there is not evidence of anything.
  • Budget time forward, not evenly. If the paper gets harder, the time you will want is at the end, which is precisely where Report 2026-01 finds candidates have none left.
  • A bad start is not a bad paper. On the Watson-Glaser the opening block is the hardest part by 28.8 points. Abandoning a paper on the strength of its first section is a real and avoidable error.
  • Practise on the shape you will sit. Rehearsing on a flat paper teaches pacing that a ramp will punish.

6A withdrawn finding this replaces

On 21 September 2026 this department published, and then withdrew within a day, the claim that time pressure degrades the answers a candidate does give. The original cut compared speeded forms against generous-clock forms and found accuracy collapsing on the former and holding on the latter.

It did not survive its own control. Splitting speeded attempts by how much of the clock they actually used, candidates who finished with time to spare declined 27.6 points against 31.3 points for the time-pressured. Almost the whole effect was present in people under no pressure at all, which means it was never about the clock.

It was a property of the paper. This report is what was left once that was taken seriously, and it is a better finding than the one it replaces: the decline is real, it is measurable, and it varies enormously between forms that candidates are told to treat as interchangeable.

7What this does not show

Position is not difficulty. A late item can be answered worse because it is harder, because the candidate is tired, or because it is being rushed. The full-coverage column removes the selection problem but not the other two, so the gradient is a property of the paper as it is experienced, which is what a candidate needs to plan for, rather than a calibrated item-difficulty curve.

  • Position is not the same as calibrated difficulty. A late item may be answered worse because it is harder, because the candidate is tired, or because it is being rushed. The full-coverage column removes selection, not fatigue.
  • These are our forms. They are written to each test’s published format and item mix, but the publishers’ live papers may order their items differently, and adaptive instruments have no fixed order at all.
  • Three forms rest on fewer than fifty attempts. Cubiks Logiks, Wonderlic and Watson-Glaser carry the widest uncertainty here, and the attempt counts are in Table 1 so a reader can weigh them.

8Sample, filters and discards

Every form with at least 30 analysed attempts, speeded and generous alike. The corpus and its filters are shared with every other report in this series.

Attempts on record1,473
Discarded: re-grade disagreed with the stored score556
Discarded: no answers recorded26
Discarded: used under half the allowed clock164
Analysed727
Distinct candidates behind those attempts575

9The forms this report is computed from

Six forms cleared the thirty-attempt floor. Each page below covers the format, and carries a free practice paper you can sit to feel the shape for yourself.

10Common questions

Do aptitude tests get harder as they go on?
Some do and some do not, and the difference is large enough to matter. On the CCAT form, accuracy falls 36.4 points between the opening fifth and the closing fifth. On the IBEW form it moves 0.1 points, which is nothing. On the Watson-Glaser form it moves the other way entirely: the opening fifth is answered at 53.1% and the closing fifth at 89.7%.
Is the CCAT harder at the end?
Yes, materially. Among CCAT attempts that answered every item, so no candidate is missing from any point on the curve, accuracy runs 81.7%, 72.8%, 73%, 56.6%, 45.3% across the five fifths. A candidate who judges their progress by how the first ten questions felt is calibrating against the easiest part of the paper.
Why is the Watson-Glaser hardest at the start?
It is a property of how that paper is ordered rather than of the candidate. The opening block is answered at 53.1% and every later block sits between 81.9% and 90.9%. The practical consequence is that a poor start on a Watson-Glaser form says much less about the eventual score than the same start on a CCAT form.
Does a ramp mean you should skip the end of the paper?
No, the opposite. A ramp means later items are answered wrong more often, not that they are worth less: every form here scores them identically and none deducts for an error. Report 2026-01 measures what actually happens to those items, which is that on a speeded form most of them are never answered at all.

Suggested citation

Aptitude tests do not share a difficulty shape. Across 6 forms, the CCAT loses 36.4 accuracy points from the opening fifth to the closing fifth and the Wonderlic 42.8, while the IBEW form is flat at -0.1 and the Watson-Glaser runs backwards at -36.6. PrepClubs Research, Report 2026-03.

Download the figures as CSVMethod and standardsAsk us to check a figure

Reproduce any figure or table with attribution to PrepClubs Research and a link to this page. The underlying per-question records are not published: joined to an attempt they are re-identifiable at this sample size.

Other reports in this series