Methodology
This page exists so that anything on this site can be checked. How the test items are built, how a score is produced, what population that score is compared against, what the test cannot do, and where every figure quoted elsewhere on the site came from. If something here is wrong, it is on this page that you should be able to catch it.
Most online IQ tests publish nothing about how they work, which is precisely why their numbers should not be trusted. The alternative is to write it down and invite the argument.
How the test items are built
Every question on this test is a matrix reasoning item: a three-by-three grid of tiles governed by hidden rules, with the bottom-right tile missing and six candidate answers offered. Your task is to work out the rules and select the tile that completes them.
Rule-based generation, not a fixed bank
Items are produced by a generator rather than drawn from a stored set of hand-drawn puzzles. The generator operates on a small vocabulary of tile attributes:
- Shape — which figure occupies the cell
- Count — how many elements appear
- Size — the scale of the elements
- Fill — solid, hollow, or patterned
- Rotation — orientation of the elements
A rule governs how one attribute changes across a row or down a column: held constant, progressing in a sequence, or combining across cells. An item is defined by which rules are active on which attributes.
How difficulty is created
Difficulty rises by layering additional rules onto the same grid, not by making the tiles harder to see. This distinction matters. A puzzle made hard through visual clutter is measuring your eyesight and patience; a puzzle made hard by requiring you to hold three interacting rules in mind at once is measuring working memory and abstract reasoning, which is what the test is for.
The thirty items are ordered by ascending difficulty. Early items carry one rule and are intended to be solvable in seconds. The final items carry several interacting rules and are genuinely hard.
The test format
| Parameter | Value |
|---|---|
| Item type | Matrix reasoning, single format throughout |
| Number of items | 30 |
| Time limit | 25 minutes, whole test |
| Answer options per item | 6 |
| Penalty for guessing | None |
| Verbal content | None |
| Registration required | No |
How scores are produced
Results are reported on the conventional deviation scale, where the population mean is set to 100 and the standard deviation to 15. This is the same metric the Wechsler scales use, which is why a score here is directly comparable in interpretation, though not in precision, to scores you may have encountered elsewhere.
From raw score to scale score
Your raw score is the number of items answered correctly out of thirty. That raw score is converted to the conventional scale — mean 100, standard deviation 15 — by a linear transformation around an assumed raw mean of 17 and an assumed raw standard deviation of 5.2. The result is rounded and clamped to a floor of 62.
Because thirty items are spread across that range, each additional correct answer is worth about 2.9 IQ points. That is the practical resolution of this test, and it is the clearest single argument for reading any result as a region rather than a point.
The highest score the test can currently return is 138, for thirty out of thirty. Scores above that are not reachable here.
What the percentile is, and what it is not
An IQ score is not a measurement of a person. It is a statement about where that person sits relative to a comparison group, and a score is only as meaningful as that group is well defined.
The percentile shown in your report is derived mathematically from your scale score, using the normal distribution. It answers: if scores were normally distributed with a mean of 100 and a standard deviation of 15, what share would fall at or below yours?
It is not derived from the people who have taken this test. No observed test-taker data enters the calculation. The scoring model is provisional: the assumed raw mean of 17 and standard deviation of 5.2 are design assumptions, not measurements.
Every completed test does store per-item responses, timings, and the raw total. Once enough completions have accumulated, the assumed parameters will be replaced with observed ones and this page will say so, with the sample size and the date. Until then, the honest description of a score here is a position on a modelled scale, not a rank among real people.
One limitation will still apply once observed norms do replace the model, and it applies to every unsupervised online test: people who choose to seek out and complete an IQ test are not a random sample of the population. They are more interested in the subject, more likely to be comfortable with abstract puzzles, and self-selected in ways that are difficult to quantify. Norms built from such a group describe a position among test-takers, not a position in the general public, and this page will say so when the time comes.
Measurement error
No cognitive test returns an exact value. Professionally administered assessments report a confidence interval around every score for this reason, and the honest reading of any single result — here or anywhere — is a region of the scale rather than a point on it.
The report currently gives a single figure rather than a range. That is a limitation of the report, not a claim to precision. Given that each correct answer moves the score by roughly 2.9 points, a difference of a few points between two sittings reflects one or two items going either way — which is ordinary variation, not a change in ability.
The factors that move a result on any given day are covered in detail on how accurate online IQ tests are. In short: fatigue, interruption, rushing and unfamiliarity with the format can each cost several points, and format familiarity from a recent previous attempt can add several. A meaningful retest should be separated by weeks.
What this test does not claim
It is not a diagnostic instrument and cannot identify a learning difficulty, an intellectual disability, giftedness, or cognitive decline. Those require formal assessment by a qualified psychologist, using evidence beyond any test score.
It is not accepted by Mensa or any other high-IQ society, all of which require supervised testing under their own conditions.
It is not valid for educational placement, employment decisions, or any legal or clinical purpose.
It measures fluid reasoning only. It makes no attempt at verbal comprehension, processing speed, or memory, which a full battery would assess separately.
How the written content is sourced
The explanatory pages on this site make empirical claims — about how ability changes with age, about what raises test scores, about how accurate online testing is. Those claims follow four rules.
- Figures come from named primary sources. Where a page states a number, that number is traceable to a specific published study, identified by author, year and DOI. Statistics are not taken from other websites, summary articles, or aggregators.
- Uncertainty is stated, not smoothed over. Where the research disagrees with itself, the page says so. Where an effect is contested or has failed to replicate, that is reported alongside the original finding rather than omitted.
- Effect sizes are described as population averages. A finding about groups is not presented as a prediction about any individual, because it is not one.
- Nothing is invented to fill a gap. If a plausible-sounding figure cannot be traced to a source, it does not go on the page. A number is either sourced or absent.
This is the standard the site holds itself to. It is also the standard by which you should judge it — if a figure appears without a traceable source, that is a failure worth reporting.
Dates, versions and corrections
Every page carries a published date and a last-updated date. Those reflect real substantive edits. Pages are not re-dated to appear fresh, and a rebuild that changes nothing does not move the date.
Corrections are made rather than quietly patched. If a factual error is found and fixed, the change is noted in the page's version history.
If you can show that something on this site is wrong — a misread statistic, a misattributed study, an item with more than one defensible answer, a scoring result that looks off — send the page and the correction to hello@startiqtest.com. It reaches the person who wrote it.
Who is responsible
This site is built and run by one person, Mathew O'Neill, through DadLink Technologies Limited (registered in England, company number 16161444). He is a software developer rather than a psychologist, and designed the item generator described above. The full statement of who is behind the test, and what that does and does not qualify him to claim, is on the about page.
Sources cited across this site
Every empirical claim on the explanatory pages traces to one of the following.
- Flynn, J. R. (1987). Massive IQ gains in 14 nations: What IQ tests really measure. Psychological Bulletin, 101(2), 171–191. doi:10.1037/0033-2909.101.2.171
- Germine, L., Nakayama, K., Duchaine, B. C., Chabris, C. F., Chatterjee, G., & Wilmer, J. B. (2012). Is the Web as good as the lab? Comparable performance from Web and lab in cognitive/perceptual experiments. Psychonomic Bulletin & Review, 19, 847–857. doi:10.3758/s13423-012-0296-9
- Hartshorne, J. K., & Germine, L. T. (2015). When does cognitive functioning peak? The asynchronous rise and fall of different cognitive abilities across the life span. Psychological Science, 26(4), 433–443. doi:10.1177/0956797614567339
- Melby-Lervåg, M., Redick, T. S., & Hulme, C. (2016). Working memory training does not improve performance on measures of intelligence or other measures of "far transfer". Perspectives on Psychological Science, 11(4), 512–534. doi:10.1177/1745691616635612
- Owen, A. M., Hampshire, A., Grahn, J. A., Stenton, R., Dajani, S., Burns, A. S., Howard, R. J., & Ballard, C. G. (2010). Putting brain training to the test. Nature, 465(7299), 775–778. doi:10.1038/nature09042
- Ritchie, S. J., & Tucker-Drob, E. M. (2018). How much does education improve intelligence? A meta-analysis. Psychological Science, 29(8), 1358–1369. doi:10.1177/0956797618774253
- Schaie, K. W. (2005). Developmental Influences on Adult Intelligence: The Seattle Longitudinal Study. Oxford University Press.
- Wechsler Adult Intelligence Scale, technical and interpretive manuals. Used for descriptive classification bands, scale conventions, and normative structure.
- Systematic reviews and meta-analyses of iodine status and child cognitive development, summarised in the review literature on iodine and mental development in children under five.
Single format, timed, price stated before you start, no registration.
Start the free IQ test30 questions. 25 minutes. No registration. Full report £9.99.