Why does my IQ score change between tests?
Last updated 1 August 2026
Because a score is an estimate, not a reading. Every cognitive test has a margin of error of several points, and repeat sittings reliably produce higher scores through familiarity alone. Two results five points apart are usually the same underlying ability measured twice.
People often treat an IQ score the way they treat a height measurement: a fixed property of themselves that a test either reads correctly or gets wrong. That framing is what makes a second, different number feel like a contradiction to be resolved. It is not. Variation between sittings is expected, quantified, and largely explainable.
Practice effects: the biggest single cause
Take any cognitive test twice and you will tend to score higher the second time. This happens without any change in ability and without any deliberate preparation.
The mechanism is unremarkable. On a second sitting you already know the instructions, you recognise the shape of the questions, you know how the timer behaves, and you are not spending the first few items working out what is being asked of you. All of that converts directly into score.
The gain is largest between the first and second attempt and diminishes after that. It is one reason clinical assessments avoid retesting the same instrument within a short window, and one reason a score from your third attempt at the same online test is the least informative number you have.
Measurement error, and the band around every score
Professionally administered tests do not report a bare number. They report a score with a confidence interval, typically several points wide in each direction, because that is the honest representation of what a test can establish in a single session.
An estimate of 112 with a band of plus or minus five means the evidence is consistent with anything from 107 to 117. If a second test returns 108, nothing has contradicted anything. Both results sit inside the same interval.
Online tests are noisier still, for reasons that have nothing to do with the questions: no supervision, no controlled environment, no check on whether you were interrupted, and no way to know whether you had done something similar the week before.
Conditions on the day
The same person tested twice under different conditions is, for practical purposes, two different test-takers.
- Sleep. Working memory is the first thing to degrade when you are tired, and matrix reasoning leans on working memory heavily.
- Time pressure. Sitting a timed test while distracted or rushed costs items at the difficult end, where the points are.
- Interruptions. A single break in concentration during a timed section can cost several questions.
- Alcohol, illness, acute stress. All measurably depress performance, none change underlying ability.
None of these raise your ceiling when removed. They stop you performing below it, which is why a rested score is a better estimate rather than a better you.
Different tests measure different things
A score from a matrix reasoning test and a score from a full clinical battery are not the same quantity, and expecting them to match is expecting too much of both.
A matrix test measures non-verbal fluid reasoning: pattern extraction with no vocabulary or general knowledge involved. A Wechsler assessment produces a Full Scale IQ from several index scores spanning verbal comprehension, perceptual reasoning, working memory and processing speed.
Most people have a spread of ten to twenty points between their strongest and weakest cognitive domains. If a narrow test happens to sample your strongest domain and a broad one averages across all of them, the two will disagree by exactly the amount of that spread. Both are correct about what they measured.
The scale a result is reported on matters too. A score of 130 on the modern SD 15 convention is 132 on SD 16 and 148 on the Cattell SD 24 scale. Identical rarity, three different numbers.
How to read two different results
The percentile is more stable than the score, because it survives rescaling. If two tests place you around the 84th percentile and the 80th, they broadly agree even if the reported numbers look different.
Treat the lower of two scores from equivalent tests as no less credible than the higher one. There is a strong pull toward believing the flattering result and explaining away the other, and no statistical basis for it.
Where two results differ substantially and it genuinely matters — for an application, an assessment, or anything with a consequence — the answer is a supervised, professionally administered test rather than a third online attempt. Online tests, including this one, estimate a position on the scale from a short question set without supervision. That is useful for curiosity and useless as evidence.
Frequently asked questions
Which of my two scores is the real one?
Is it normal for an IQ score to go up on a retest?
How much can a score vary between sittings?
Does a lower second score mean I have got worse?
More in this series
- Does brain training increase IQ?Near transfer, far transfer, and why the games get easier while nothing else does.
- Does dual n-back work?The 2008 result that started it, and what the placebo-controlled replication found.
- Does education increase IQ?The one intervention with strong evidence: 1 to 5 points per additional year.