Methodology

How the test is built and scored

Most online IQ tests never explain how they reach your number. This page shows exactly how the IQScore assessment is put together, the precise formula that turns your answers into a score, and an honest account of what that score can and cannot tell you.

The instrument

The assessment is 36 multiple-choice items that we wrote and calibrated ourselves. Every item is a reasoning problem with a single correct answer and no knowledge prerequisite, so the score reflects reasoning ability rather than education or vocabulary.

The items are spread across four cognitive domains and arranged on a rising difficulty curve, so the test stays informative for both average and high-ability test-takers:

Spatial reasoning

Matrix and pattern-completion items: identifying the rule that governs a visual sequence and selecting the figure that completes it.

Numerical reasoning

Number series and quantitative relationships that test inductive reasoning without requiring advanced maths.

Logical reasoning

Deductive and inferential problems where the answer follows from stated premises.

Applied reasoning

Problems that combine the above under realistic constraints, closer to how reasoning is used in practice.

What the test measures

The four domains are all indicators of general cognitive ability, the factor psychologists label g. They map onto the Cattell-Horn-Carroll (CHC) model of intelligence, the framework that underlies most modern professional assessments such as the WAIS and the Stanford-Binet. Fluid reasoning (Gf) carries the most weight, which is why pattern and matrix items make up the largest share of the test.

What the test does not measure is just as important. It does not assess working memory, processing speed, personality, motivation, or emotional intelligence. A single reasoning score is a useful signal, not a complete picture of a mind.

How we score

Scoring is deterministic and difficulty-weighted. Harder items are worth more points, so two people who answer the same number of questions correctly can receive different scores depending on which ones they got right. The steps are:

1. Sum the points of every correct answer  →  raw score (0–72)
2. Locate the raw score in the band table below
3. Interpolate linearly within the band to a single IQ value
iq = bandMinIQ + (raw − bandMinPts) / (bandMaxPts − bandMinPts) × (bandMaxIQ − bandMinIQ)

The bands are anchored to the standard IQ metric, a mean of 100 and a standard deviation of 15, and clamped to a 55–145 range. This is the full mapping, taken directly from the live scoring code:

Raw pointsIQ rangeClassificationPercentile
65–72130–145Very Superior98th+
55–64120–129Superior91st–97th
43–54110–119High Average75th–90th
30–4290–109Average25th–74th
20–2980–89Low Average9th–24th
10–1970–79Below Average2nd–8th
0–955–69Well Below AverageBelow 2nd

Your results page also shows a per-domain breakdown, the share of available points you earned in each of the four domains, so a single number never hides the shape of your performance.

What we have validated, and what we haven't

This is where most test sites go quiet or borrow someone else's numbers. We won't do either, so here is the plain truth.

What stands behind the score. The score bands are anchored to the same mean-100, SD-15 metric used by clinical instruments, and the difficulty ordering of the items is checked against real response data. We have collected more than 11,000 completed assessments, which gives us item-level statistics we use to refine difficulty and tighten the bands over time.

What we have not done. That 11,000-person dataset is a self-selected online sample, not a nationally representative norming panel. We have not run an independent standardisation study, and we do not publish reliability or validity coefficients we cannot stand behind. Anchoring to the standard IQ scale is not the same as clinical norming, and we won't pretend otherwise.

Where this is heading. We are using the growing response dataset to validate our own items properly: item-difficulty and discrimination analysis first, then internal-consistency estimates. As those numbers become defensible, they will be published here, with their limitations stated alongside them.

Known limitations

It is a screening estimate.

A 25-minute online test produces an estimate with a genuine margin of error, not a diagnostic score. Treat a result as a range, not a precise figure.

Conditions are not controlled.

We cannot verify your environment, effort, or whether you were interrupted. Distraction, fatigue, or a second attempt all move the number.

Practice and exposure inflate scores.

Reasoning-test performance improves with familiarity. A score after several online tests reflects practice as much as ability.

It is a snapshot of reasoning only.

The test captures fluid reasoning at one moment. It says nothing about memory, speed, knowledge, or the many other things intelligence involves.

How this compares to a clinical test

A professionally administered test such as the WAIS-V or Stanford-Binet 5 is delivered one-to-one by a licensed psychologist, under controlled conditions, and scored against a representative norming sample. That is the right standard for any clinical, educational, or employment decision.

IQScore is not that, and does not try to be. It is a free, instant, self-administered estimate for curiosity and self-insight. Used that way, and read with its limitations in mind, it is genuinely useful. Used as a diagnosis, it would be misused.

See where you rank

Free · 36 questions · Instant results · No sign-up

Take the Free IQ Test →About IQScore