Methodology
How the test is built and scored
Most online IQ tests never explain how they reach your number. This page shows exactly how the IQScore assessment is put together, the precise formula that turns your answers into a score, and an honest account of what that score can and cannot tell you.
The instrument
The assessment is 36 multiple-choice items that we wrote and calibrated ourselves. Every item is a reasoning problem with a single correct answer and no knowledge prerequisite, so the score reflects reasoning ability rather than education or vocabulary.
The items are spread across four cognitive domains and arranged on a rising difficulty curve, so the test stays informative for both average and high-ability test-takers:
Spatial reasoning
Matrix and pattern-completion items: identifying the rule that governs a visual sequence and selecting the figure that completes it.
Numerical reasoning
Number series and quantitative relationships that test inductive reasoning without requiring advanced maths.
Logical reasoning
Deductive and inferential problems where the answer follows from stated premises.
Applied reasoning
Problems that combine the above under realistic constraints, closer to how reasoning is used in practice.
What the test measures
The four domains are all indicators of general cognitive ability, the factor psychologists label g. They map onto the Cattell-Horn-Carroll (CHC) model of intelligence, the framework that underlies most modern professional assessments such as the WAIS and the Stanford-Binet. Fluid reasoning (Gf) carries the most weight, which is why pattern and matrix items make up the largest share of the test.
What the test does not measure is just as important. It does not assess working memory, processing speed, personality, motivation, or emotional intelligence. A single reasoning score is a useful signal, not a complete picture of a mind.
How we score
Scoring is deterministic and difficulty-weighted. Harder items are worth more points, so two people who answer the same number of questions correctly can receive different scores depending on which ones they got right. The steps are:
The bands are anchored to the standard IQ metric, a mean of 100 and a standard deviation of 15, and clamped to a 55–145 range. This is the full mapping, taken directly from the live scoring code:
| Raw points | IQ range | Classification | Percentile |
|---|---|---|---|
| 65–72 | 130–145 | Very Superior | 98th+ |
| 55–64 | 120–129 | Superior | 91st–97th |
| 43–54 | 110–119 | High Average | 75th–90th |
| 30–42 | 90–109 | Average | 25th–74th |
| 20–29 | 80–89 | Low Average | 9th–24th |
| 10–19 | 70–79 | Below Average | 2nd–8th |
| 0–9 | 55–69 | Well Below Average | Below 2nd |
Your results page also shows a per-domain breakdown, the share of available points you earned in each of the four domains, so a single number never hides the shape of your performance.
What we have validated, and what we haven't
This is where most test sites go quiet or borrow someone else's numbers. We won't do either, so here is the plain truth.
What stands behind the score. The score bands are anchored to the same mean-100, SD-15 metric used by clinical instruments, and the difficulty ordering of the items is checked against real response data. We have collected more than 11,000 completed assessments, which gives us item-level statistics we use to refine difficulty and tighten the bands over time.
What we have not done. That 11,000-person dataset is a self-selected online sample, not a nationally representative norming panel. We have not run an independent standardisation study, and we do not publish reliability or validity coefficients we cannot stand behind. Anchoring to the standard IQ scale is not the same as clinical norming, and we won't pretend otherwise.
Where this is heading. We are using the growing response dataset to validate our own items properly: item-difficulty and discrimination analysis first, then internal-consistency estimates. As those numbers become defensible, they will be published here, with their limitations stated alongside them.
Known limitations
It is a screening estimate.
A 25-minute online test produces an estimate with a genuine margin of error, not a diagnostic score. Treat a result as a range, not a precise figure.
Conditions are not controlled.
We cannot verify your environment, effort, or whether you were interrupted. Distraction, fatigue, or a second attempt all move the number.
Practice and exposure inflate scores.
Reasoning-test performance improves with familiarity. A score after several online tests reflects practice as much as ability.
It is a snapshot of reasoning only.
The test captures fluid reasoning at one moment. It says nothing about memory, speed, knowledge, or the many other things intelligence involves.
How this compares to a clinical test
A professionally administered test such as the WAIS-V or Stanford-Binet 5 is delivered one-to-one by a licensed psychologist, under controlled conditions, and scored against a representative norming sample. That is the right standard for any clinical, educational, or employment decision.
IQScore is not that, and does not try to be. It is a free, instant, self-administered estimate for curiosity and self-insight. Used that way, and read with its limitations in mind, it is genuinely useful. Used as a diagnosis, it would be misused.
See where you rank
Free · 36 questions · Instant results · No sign-up