PrepClubs Research

How we measure things

Everything we publish comes from practice attempts on this platform. That is a real limitation and we would rather state it than bury it: these are people preparing for a test, not people sitting one for an employer, so the population is self-selected and skews motivated.

Two filters

One, we re-grade everything. Each attempt is re-scored from its stored answers against the question bank and kept only when the recomputed score matches the stored one. Our banks have been edited since launch, so an older attempt can no longer be scored against the questions it was actually sat on. This is the filter that discards the most.

Two, did they actually sit it. We keep only attempts that used at least half the allowed time, which removes people who opened a test and wandered off.

Half and not ninety per cent, and the reason matters. A ninety per cent threshold selects for time pressure, so it quietly removes the people who found a test easy and handed in early. Applied to a comparison between a fast test and a slow one, it would thin the slow group of exactly the candidates that make it the control. We report the ninety per cent cut only as a robustness check.

What we do not filter on: whether somebody finished. On a study about running out of time, not finishing is the finding. Excluding those attempts would mean selecting on the outcome being measured, which would make the result meaningless and flattering at the same time.

The discard ledger

Attempts on record1438
Re-grade disagreed555
Empty25
Used under half the clock161
Analysed697

What we cannot tell you

  • Anything by country, device or demographic. We do not collect it.
  • Anything about percentiles or national norms. We have no norm table and will not imply one.
  • Anything about IT certification practice. Seven usable attempts is not a sample.
  • Whether a second sitting helps. Of eighty repeat pairs, sixty-one re-sat the same test, so the improvement is largely item memory. We are still measuring it. We are not publishing it, and we would rather say that than publish a number that flatters us.

Withdrawn findings

“Rushing degrades the answers you do give”, withdrawn 21 September 2026

An early cut found accuracy on answered questions falling from 83.4% to 52.3% across a speeded test while staying flat on a generous one, and we read that as the clock degrading judgement. It does not hold. That comparison confounds the clock with test design: the CCAT ramps difficulty and the Watson-Glaser does not.

Tested properly, within speeded tests only and split by whether the candidate was actually short of time, those who finished with time to spare declined 27.6 points against 31.3 for those who did not. Almost all of the fall is a difficulty ramp. Our banks carry no difficulty labels, so this split is the only way to separate the two, and any future claim about accuracy by position has to clear it first.

We do not delete findings. A withdrawn result keeps its place here with the reason, because a study you cannot check is a press release.

Ask us to check something

If you want a figure verified, or a cut we have not published, ask and we will run it. If the answer is unflattering we will publish that too. Get in touch.