Skip to content
AnamneseArchive of Psychology

At least two characters

The Journal14 December 20259 min read

What the IQ measures is not what it promises

The Flynn effect, heritability, and the superlative that fell in 2022

About three points per decade across the twentieth century — and since the 1990s the reversal, visible in the Norwegian data even between brothers of the same parents.Drawing by the archive

For four decades a single figure carried the career of the intelligence test. In 1998 Frank L. Schmidt and John E. Hunter summarised what hundreds of selection studies had produced on the relation between test performance and later occupational success, and arrived at a correlation of r = .51. No other selection procedure reached that value — not the job interview, not the reference, not the work sample. The number travelled into the textbooks, into expert reports, into personnel departments.

In 2022 Paul R. Sackett, Charlene Zhang, Christopher M. Berry and Filip Lievens recalculated it. Their finding concerns not the data but a computational convention: the usual correction for range restriction — the statistical extrapolation of a relation measured only on applicants who had already been pre-selected — systematically overcorrects. The new estimate for occupational success stands at ρ = .31, and the structured interview moves ahead of the intelligence test. The superlative the field had tended for a century is not thereby refuted. It is struck out.

What remains is a double balance sheet of a kind no other instrument in this archive carries. On one side: IQ is still among the strongest single predictive measures psychology has ever built. School success, occupational success, even life expectancy correlate with it, and more reliably than with most of the field’s constructs. On the other side: no measure has been more thoroughly misused — as scientific costume for immigration quotas, forced sterilisations and racial hierarchies. Whoever tells only one of the two balances tells it wrong.

What the test was originally built for

The birth was solicitous. In 1904 the French ministry of education was looking for a procedure to identify children who could not follow regular schooling and needed support. In 1905 Alfred Binet and Théodore Simon built a series of tasks of rising difficulty and compared the individual child with its age-mates. Binet himself warned against both of the things that came later: the measure was no fixed organ, intelligence not a single number, and whoever read the value as fate was misusing it. He died on 18 October 1911, before he had to watch how right he would turn out to be.

He was not even the first. From 1895 Hermann Ebbinghaus worked for a commission of the Breslau school authority, and in 1897 he presented his combination or completion method: gapped texts the child had to close sensibly. He had built it to measure mental fatigue across the school day; what made it useful was its value as a measure of general capability — eight years before Binet and Simon. This archive’s Ebbinghaus file tells the case in detail. It shows that the measurement of intelligence did not grow out of a theory of intelligence, but out of the administrative needs of school authorities.

How care turned into selection

In America the instrument for finding support needs became a sorting machine. Lewis Terman normed the test at Stanford, believed in the heritability of the measured value and in its fitness for selection. Between 1917 and 1919 Robert Yerkes tested roughly 1.75 million recruits of the US Army, about 1.5 million of them with the written Army Alpha. The results were methodologically worthless and politically highly effective: command of English and years of schooling entered the scores without being separated out, a textbook case of confounding. A man who had been in the country for three years looked, on paper, like a feeble-minded one.

Those data became ammunition. Immigration quotas against southern and eastern Europeans invoked them; sterilisation laws came into force in around 32 states, upheld by the Supreme Court in 1927 in Buck v. Bell. The wording belongs here, because any paraphrase softens the coldness:

The principle that sustains compulsory vaccination is broad enough to cover cutting the Fallopian tubes. … Three generations of imbeciles are enough.

Oliver Wendell Holmes Jr., Buck v. Bell, 274 U.S. 200 (1927)

The number of people forcibly sterilised in the United States runs to the tens of thousands. It is this archive’s best-documented case of a psychological measure acquiring the force of law and destroying human lives. No error of the test was to blame. What was to blame was the readiness to read a group mean as a rank of nature.

What the prediction carries and what it does not

Post-war research surveyed the object more soberly. The general factor — that performances across the most varied tasks correlate positively — is among the field’s most stable findings; the quarrel is not about the correlation but about its interpretation. The predictive power for school and work is real and secured in meta-analyses, though its magnitude is under revision right now, as the recalculation by Sackett and colleagues shows. Heritability within populations is likewise considerable, and it rises with age. This text deliberately gives no point estimate, because the values scatter widely by age and sample.

And exactly here sits the quarrel’s most consequential fallacy: heritability within a group explains no differences between groups. Height is highly heritable — and a generation with better nutrition grows by centimetres anyway. The same argument holds for group differences in IQ, and the environmental levers have names: years of schooling, nutrition, lead, test familiarity, poverty.

Part of the sobriety is the balance of interventions, and it falls in two. Early programmes such as Head Start raise measured scores markedly at first; part of the gain fades across the school years — the notorious fade-out. At the same time, long-run data show stable effects where it counts. David Deming found effects of Head Start on school completion and further education. Jens Ludwig and Douglas Miller, comparing across the funding threshold, found effects reaching into educational attainment and mortality. The test score fades; the life course does not.

The evidence on schooling is clearer. In 2018 Stuart Ritchie and Elliot Tucker-Drob assembled what quasi-experimental studies had produced: 142 effect sizes from 42 datasets with over 600,000 participants, resting on three independent families of design — school entry cut-off dates, compulsory-schooling reforms and controls for prior ability. Each additional year of school raises measured IQ by roughly one to five points. What the test measures is trainable. Just not with miracle cures, but with years.

What the Flynn effect rules out

The strongest evidence against any fixed reading bears a name: James Flynn. Across the twentieth century, raw scores rose in practically every country studied by around three points per decade — so fast that genes drop out as an explanation. A generation measured against its grandparents' norms would sit in the gifted range. Nobody believes that is real, so the test is also measuring familiarity with abstract formats of thought that modern societies train into people.

Since the 1990s the curves have stalled or fallen again in several rich countries. The interpretations range from school systems to measurement artefacts, but one explanation can be excluded. Bernt Bratsberg and Ole Rogeberg analysed the Norwegian conscription data and found the decline within families — that is, between brothers who have the same parents. What changes between two brothers is not genetics. A measure whose mean shifts by standard deviations within a few generations is measuring something real and nothing fixed. The Flynn effect is therefore not a curiosity at the edge of the quarrel but its single most important finding.

StationYearWhat happenedWhat followed from it
Ebbinghaus, completion method1897gap texts for the Breslau school authoritymeasurement came before theory
Binet & Simon1905a series of tasks to identify need for supportthe child is compared with its age group
Yerkes, army tests1917–1919around 1.75 million recruitslanguage and schooling enter the score unseparated
Buck v. Bell1927the Supreme Court upholds forced sterilisationlaws in around 32 states
Flynn198714 countriesa rise of about three points per decade
Ritchie & Tucker-Drob2018142 effect sizes, over 600,000 participantsone school year raises scores by one to five points
Bratsberg & Rogeberg2018Norwegian conscription datathe reversal appears between brothers as well
Sackett et al.2022re-estimate of the validityr = .51 becomes ρ = .31

How the field argues with itself

Two chapters are regularly cut short. The first is the Bell Curve debate of 1994: Richard Herrnstein and Charles Murray tied IQ to class and origin and set off the largest public exchange in the field’s history. The American Psychological Association’s answer came as a consensus report: “Intelligence: Knowns and Unknowns”, written under the chairmanship of Ulric Neisser, whose file lies in this archive. The report remains the model for how a field can answer a politicised sharpening. Take stock, separate knowledge from ignorance, mark the between-group causal question explicitly as open.

The second chapter is test fairness. Out of the justified critique of culture-laden items grew a methodology of its own: measurement invariance testing, culture-reduced matrix tests, adaptive procedures. Today’s instruments are not the army tests of 1917. Whoever cites the misuse of then against the measurement of now makes the same error in the opposite direction.

How hard self-examination is, is shown by its most famous document of all. Stephen Jay Gould’s The Mismeasure of Man accused the skull measurer Samuel George Morton of unconscious manipulation. In 2011 Jason Lewis and colleagues re-measured Morton’s collection and reversed the charge: it was not Morton but Gould who had calculated with a bias. That counter-critique has in turn been disputed, by Michael Weisberg in 2014 and by Jonathan Kaplan and colleagues in 2015. The quarrel is open, and it is instructive: the exposure of bias is not itself free of bias.

What remains

The test measures something stable and predictive, but only within one society and at one time. How strong the prediction is has been under renegotiation since 2022, and on the current state of knowledge it comes out lower than was taught for decades. What the test does not measure is a person’s worth, and it is no fixed natural endowment: it can be moved by years of schooling, and its mean has migrated by standard deviations across the twentieth century. Group means say nothing about individuals and little about causes — that is the one sentence to take away from this whole quarrel.

In everyday life the misuse is recognisable by a simple sign: a number is read as a property rather than as a result measured under conditions. A test score is always a score of this person, on this day, with this instrument, against this norming sample. Drop any one of those four specifications and the number becomes a label. Binet’s invention was meant to get children support. It took a century to bring it back there — and the most dangerous property of a number remains its apparent neutrality.

Sources, and why they are here

  1. Ebbinghaus, H. (1897). Über eine neue Methode zur Prüfung geistiger Fähigkeiten und ihre Anwendung bei Schulkindern. Zeitschrift für Psychologie und Physiologie der Sinnesorgane, 13, 401–459.

    The German-language prehistory, eight years before Binet: intelligence measurement arose not from a theory of intelligence but from the administrative need of a school authority.

  2. Binet, A., & Simon, T. (1905). Méthodes nouvelles pour le diagnostic du niveau intellectuel des anormaux. L'Année Psychologique, 11, 191–244.

    The hour of birth — and it was meant kindly. Binet himself warned against both of the things it became: the fixed number and the verdict of fate.

  3. Buck v. Bell, 274 U.S. 200 (1927).

    The low point, and the source of the quotation. It stands in its own words because paraphrase softens the coldness of the sentence, and that coldness is part of the matter.

  4. Flynn, J. R. (1987). Massive IQ gains in 14 nations: What IQ tests really measure. Psychological Bulletin, 101, 171–191.

    The strongest single finding against any fixed reading: about three points per decade, too fast for genes.

  5. Bratsberg, B., & Rogeberg, O. (2018). Flynn effect and its reversal are environmentally caused. PNAS, 115, 6674–6678.

    Rules out an explanation of the reversal instead of merely offering interpretations: the decline shows up between brothers of the same parents.

  6. Neisser, U., Boodoo, G., Bouchard, T. J., et al. (1996). Intelligence: Knowns and unknowns. American Psychologist, 51, 77–101.

    The consensus report after the Bell Curve debate — the model for how a field answers a politicised provocation: take stock, name what is not known.

  7. Ritchie, S. J., & Tucker-Drob, E. M. (2018). How much does education improve intelligence? A meta-analysis. Psychological Science, 29(8), 1358–1369.

    The clearest evidence in the whole question: 142 effect sizes, more than 600,000 participants, three independent families of design. What the test measures is trainable — over years.

  8. Deming, D. (2009). Early childhood intervention and life-cycle skill development: Evidence from Head Start. American Economic Journal: Applied Economics, 1(3), 111–134.

    One half of the intervention balance sheet: the test score fades, the effect on educational attainment does not.

  9. Ludwig, J., & Miller, D. L. (2007). Does Head Start improve children's life chances? Evidence from a regression discontinuity design. Quarterly Journal of Economics, 122(1), 159–208.

    The other half, by way of a comparison at the funding threshold — with effects reaching into educational attainment and mortality.

  10. Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology. Psychological Bulletin, 124, 262–274.

    The figure that carried the test's career for four decades: r = .51, unmatched by any other selection method.

  11. Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.

    The correction the article opens with. It was not the data that were wrong but a computational convention — and afterwards the structured interview stands first.

  12. Gould, S. J. (1981/1996). The Mismeasure of Man. Norton.

    The most famous document of self-examination — and itself become the object of one.

  13. Lewis, J. E., DeGusta, D., Meyer, M. R., Monge, J. M., Mann, A. E., & Holloway, R. L. (2011). The mismeasure of science: Stephen Jay Gould versus Samuel George Morton on skulls and bias. PLoS Biology, 9(6), e1001071.

    The reversal of the charge after the collection was remeasured: not Morton but Gould, it argued, had calculated with a bias.

  14. Weisberg, M. (2014). Remeasuring man. Evolution & Development, 16(3), 166–178; and Kaplan, J. M., Pigliucci, M., & Banta, J. A. (2015). Gould on Morton, redux. Studies in History and Philosophy of Biological and Biomedical Sciences, 52, 22–31.

    The counter-criticism of that reversal. The quarrel stays open — and teaches that even the exposure of bias is not free of bias.