What a test score means, and how to convert it
Enter a score on any common scale and see it on all the others, with the interpretation that should travel alongside it.
Score converter
Enter one value. The rest are worked out from it, assuming a normal distribution.
The scales, and why there are so many
Every one of these is the same underlying position expressed differently. The z-score is the raw statement: how many standard deviations above or below the mean a person sits. The others exist to avoid negative numbers, decimals or false precision when a result is written into a report or read out to someone.
| Scale | Mean | SD | Where you meet it |
|---|---|---|---|
| z-score | 0 | 1 | Research papers, statistical work |
| T-score | 50 | 10 | Behaviour rating scales, personality inventories |
| Scaled score | 10 | 3 | Subtests of Wechsler batteries |
| Standard score | 100 | 15 | Composite and index scores, IQ-type indices |
| Percentile | 50 | — | Feedback to non-specialists |
| Stanine | 5 | 2 | Educational testing |
Percentiles are not evenly spaced
This is the mistake that causes the most misreading. Moving from the 50th to the 60th percentile is a quarter of a standard deviation. Moving from the 88th to the 98th is more than a full one. Because scores bunch up in the middle of a normal distribution, small differences near the centre look dramatic as percentiles and large differences at the edges look modest. When you need to compare two results, compare them in standard deviations, then translate to percentiles only for the explanation.
How wide is the band around a score?
A reported score is an estimate, not a reading off a ruler. Its confidence interval depends on the reliability of the instrument. With a reliability of 0.90 and a standard score scale, the ninety-five per cent interval is roughly nine points either side, so a score of 104 and a score of 96 are not meaningfully different. Any report that gives a single number with no interval is telling you less than it appears to.
Two things a score cannot do
- It cannot tell you why performance was what it was. Sleep, medication, anxiety, effort, language and unfamiliarity with testing all move scores, and none of them are the ability being measured.
- It cannot be compared across instruments with different normative samples. Two tests can both report a standard score of 95 and mean genuinely different things if one was normed on a national sample and the other on undergraduates.
How the CogniFit scale works
The platform reports on a 0 to 800 scale, where 800 is the maximum. Results are broken into five broad domains, reasoning, memory, attention, motor and perception, and further into twenty-three individual skills such as inhibition, shifting, updating, working memory, planning and response time. That structure is what lets a single session produce both a headline figure and a profile showing which components carried it.
The exercises named on this page are part of the CogniFit platform and are used here with permission. They are cognitive stimulation activities for general use. They are not a medical device, they do not diagnose or treat any condition, and a score from any of them is not a clinical finding. If you are worried about your memory or concentration, speak to a health professional.
Common questions
How do I convert a z-score to a percentile?
Find the area under the normal curve to the left of the z-score. A z of 0 is the 50th percentile, +1 is roughly the 84th and -1 roughly the 16th. The converter above does it for any value.
What is a T-score?
A rescaled z-score with a mean of 50 and a standard deviation of 10, used widely on behaviour rating scales because it avoids negative numbers and decimals.
Why is the difference between the 50th and 60th percentile smaller than it looks?
Because scores bunch up in the middle of a normal distribution. Ten percentile points near the centre is about a quarter of a standard deviation; ten points near the top can be more than a full one.
Can I compare scores from two different tests?
Only cautiously. Two tests reporting the same standard score can mean different things if they were normed on different samples. Comparison is safest within a single battery normed on one population.