How IQ Tests Work – Scoring, Questions & Methodology
IQ tests may seem mysterious, but they follow well-established principles of psychometric science. Behind every IQ score is a carefully structured process that involves standardized administration, specific question types, statistical norming, and conversion formulas. This guide explains how modern IQ tests work from start to finish, so you can understand exactly what happens between answering the first question and receiving your score.
The Basics of IQ Test Administration
Professionally administered IQ tests are given under controlled conditions by a trained psychologist or psychometrician. The examiner follows a strict protocol, reading instructions verbatim, timing sections precisely, and recording responses according to a standardized scoring manual. This consistency is crucial because IQ scores are only meaningful when everyone takes the test under the same conditions.
Most clinical IQ tests are administered one-on-one in a quiet room, with the examiner sitting across from the test-taker. Sessions typically last between 60 and 90 minutes for adults, though comprehensive batteries can run longer. The examiner may use physical materials (blocks, puzzle pieces, picture cards) in addition to verbal questions and written booklets.
Online IQ tests, including the free test on this site, cannot replicate these controlled conditions perfectly. However, they can still provide a useful estimate by using well-designed questions and statistical scoring. For a deeper comparison, see our article on online versus official IQ tests.
Types of Subtests and Question Formats
Modern IQ tests are not a single monolithic exam. They consist of multiple subtests, each measuring a different cognitive ability. The Wechsler Adult Intelligence Scale (WAIS), for example, includes 10 core subtests organized into four index scores:
- Verbal Comprehension Index – Subtests like Similarities (explaining how two concepts are alike), Vocabulary (defining words), and Information (answering general knowledge questions). These measure crystallized intelligence, which reflects accumulated knowledge and verbal reasoning.
- Perceptual Reasoning Index – Subtests like Block Design (replicating patterns with colored blocks), Matrix Reasoning (completing visual patterns), and Visual Puzzles. These measure fluid reasoning and spatial processing.
- Working Memory Index – Subtests like Digit Span (repeating number sequences forward, backward, and in ascending order) and Arithmetic (solving math problems mentally). These assess the ability to hold and manipulate information in short-term memory.
- Processing Speed Index – Subtests like Symbol Search and Coding, where you quickly scan and match symbols under time pressure. These measure cognitive efficiency and attention.
Other tests emphasize different formats. Raven's Progressive Matrices uses only non-verbal pattern completion, while the Stanford-Binet 5 includes both verbal and non-verbal routing subtests along with measures of fluid reasoning, knowledge, quantitative reasoning, visual-spatial processing, and working memory.
How Raw Scores Are Converted to IQ Scores
When you finish an IQ test, the examiner first calculates your raw score on each subtest – essentially, the number of items you answered correctly (sometimes with partial credit). But raw scores are not IQ scores. A raw score of 25 out of 30 on one subtest might represent very different levels of ability depending on how difficult the questions were and how the general population performed.
To make scores comparable, raw scores are converted to scaled scores using a norming table. Each subtest's raw score is mapped to a scaled score with a mean of 10 and a standard deviation of 3. This means a scaled score of 10 on any subtest represents exactly average performance for your age group.
The scaled scores from related subtests are then combined into index scores (such as the Verbal Comprehension Index), and finally, all index scores are combined to produce the Full-Scale IQ (FSIQ). Both index scores and the FSIQ use a scale with a mean of 100 and a standard deviation of 15. For the detailed mathematics of this process, see our guide on how IQ is calculated.
The Role of Standard Deviation
Standard deviation is one of the most important concepts in understanding IQ scores. On the most widely used IQ scales (Wechsler and Stanford-Binet 5), the standard deviation is set at 15 points. This means:
- About 68% of the population scores within one standard deviation of the mean (between 85 and 115).
- About 95% scores within two standard deviations (between 70 and 130).
- About 99.7% scores within three standard deviations (between 55 and 145).
Some older tests, like the original Cattell scales, used a standard deviation of 24 points instead of 15. This means a Cattell IQ of 148 is roughly equivalent to a Wechsler IQ of 130 – both represent the same percentile rank (approximately the 98th percentile). Always check which scale a test uses before comparing scores across different assessments.
Confidence Intervals and Measurement Error
No psychological test produces perfectly precise measurements. Every IQ score comes with a confidence interval – a range of scores within which your true ability most likely falls. For example, if your measured IQ is 112, the 95% confidence interval might be 107 to 117. This means there is a 95% probability that your true IQ lies somewhere within that range.
The width of the confidence interval depends on the test's standard error of measurement (SEM), which is influenced by the test's reliability. The WAIS-IV Full-Scale IQ, for instance, has a reliability coefficient of about 0.98 and an SEM of approximately 2.16 points, making it one of the most precise psychological instruments available. Online tests typically have larger confidence intervals because they use fewer questions and lack controlled administration conditions.
This is why psychologists emphasize that an IQ score is best understood as an estimate within a range rather than a single fixed number. For more on this topic, see our article on IQ test accuracy and limitations.
Standardization and Norming
The reliability of an IQ test depends heavily on its norming sample. When a new edition of an IQ test is developed, the publisher administers it to thousands of people who are carefully selected to represent the broader population in terms of age, gender, education, ethnicity, and geographic region.
For example, the WAIS-IV was normed on 2,200 adults aged 16 to 90 across the United States. The sample was stratified to match U.S. Census data on key demographic variables. This norming process is what allows the test to produce meaningful scores – your performance is being compared to a known, representative reference group.
Norming samples are updated with each new edition of a test, which is one reason why tests are revised every 10 to 20 years. The Flynn effect – the observed rise in average IQ scores over time – means that norming data must be refreshed regularly to keep the average at 100.
Ready to Test Your IQ?
Take our free 30-question IQ test and get your estimated score instantly.
Take the Free IQ Test