The Journal7 June 20268 min read
The Barnum effect: why horoscopes work
Thirteen sentences, 39 students, a mean of 4.26
In 1948 the psychologist Bertram Forer had 39 students fill in a personality test and handed out the individual evaluations a week later — thirteen sentences, with the request to rate their accuracy on a scale from zero to five. The mean was 4.26. Only once the ratings had been collected did Forer reveal what he had done: the sheet was the same for everyone, assembled sentence by sentence from a newsstand astrology book. The test itself was never scored. “You have a great need for other people to like you.” “At times you are extraverted and sociable, at times wary and reserved.” “You have a great deal of unused capacity.”
When the inferences are universally valid, as they often are, the confirmation is useless.
Forer called the process the fallacy of personal validation: people do not check a description of themselves against whether it also fits others — only against whether their own life supplies confirming instances. And it always does, if the sentences are built properly. In 1956 Paul Meehl gave the phenomenon the name under which it became famous: the Barnum effect. He took the phrase from his colleague Donald G. Paterson, who had spoken of personality descriptions “of the P. T. Barnum type” — after the showman credited with the principle of a little something for everybody. Meehl coined the term in a paper meant to persuade clinicians to adopt statistical rules of evaluation, and he meant it as a warning sign: a finding the described person agrees with is not thereby a finding about her.
| Rating from zero to five | 0 | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|---|
| The test itself | 0 | 0 | 0 | 1 | 25 | 13 |
| The evaluation handed out | 0 | 0 | 1 | 4 | 18 | 16 |
The construction manual for an apt description
The research after Forer dissected the ingredients, and they can be learned. First, two-sidedness: “sometimes this, sometimes that” sentences cannot miss, because they cover both halves of the spectrum — and still feel like an observation, because everyone knows both halves. Second, flattery with a caveat: unused capacities, hidden depth, self-criticism as a mark of high standards. People accept positive statements about themselves more readily, and most readily those that imitate modesty. Third, the illusion of personalisation: identical sentences are rated more accurate when they were supposedly produced for the reader alone — the mere frame “your evaluation” raises assent, as studies with identical texts under different labels show. And fourth, the authority of the source: a finding from an expert, from a computer or from an old system with a terminology of its own strikes home more precisely, allegedly, than the same finding from corridor gossip.
Two reviews have surveyed the field — Dickson and Kelly in 1985, Furnham and Schofield in 1987 — and reach the same result: of the four ingredients, flattery is the most dependable. Unfavourable descriptions are accepted with markedly more reserve, and that holds even when they arrive with the same authority.
That describes the market built on the effect. Horoscopes are its oldest application; graphology and palmistry work from the same kit. But the most wounding point of the research concerns psychology itself. Type tests that sort people into a few uniformly benevolently described kinds produce the same assent by the same route — the descriptions are general, positive and two-sided, and the readers validate them personally. This archive’s Jung file tells how a conceptual order became a billion-dollar business with sixteen kinds of person; the Barnum effect explains why the customers stay satisfied although a substantial share are assigned a different type on retest. Satisfaction is not a criterion of quality. It is the product.
What the matching test shows
Assent and accuracy are two different quantities, and they can be pulled apart. The procedure for doing so is simple: instead of asking how well one interpretation fits, you present several and have the right one picked out. Anyone who genuinely describes has to hit the difference; anyone producing Barnum sentences cannot, because all the interpretations fit everybody.
That is exactly the test the physicist Shawn Carlson ran on astrology in Nature in 1985, double-blind and with procedural rules agreed in advance with professional astrologers. Astrologers nominated by their own associations as particularly qualified each received a birth chart and three personality profiles from the California Psychological Inventory, and were asked to say which profile belonged to the chart. They did not hit the target more often than chance. In a second part, test subjects were given three interpretations, one of which came from their own birth chart — they too failed to find their own at above-chance rates. The same material that looks astonishingly precise in a single case loses its precision the moment it has to compete.
What makes a statement testable
The effect does not render every personality description worthless, and this is exactly where it becomes instructive. A usable statement about a person differs from a Barnum statement in three properties. It is falsifiable — it is conceivable that it does not apply to someone. It discriminates — it fits some people markedly better than others, which is what the matching test checks. And it predicts something that is not already contained in the question. Serious personality assessment has imposed these tests on itself and calls the result validity; this archive’s Allport file describes how eighteen thousand words about people ended up as five dimensions whose scores predict school grades, occupational success and health modestly but demonstrably. The difference between a factor score and a horoscope is not the tone — it is the possibility of being wrong.
A side glance at the everyday life of science itself is worth taking, because the effect does not stop at the laboratory door. Feedback from assessments, strengths profiles from HR departments, the result reports of some online studies — wherever people receive individual evaluations assembled from text modules, Forer’s seminar of 1948 runs in miniature. The test is always the same and takes a minute: swap sheets with your neighbour. Descriptions that survive the swap unharmed say something about the modules; only those the swap ruins said something about the person.
Why the effect does not go away
That leaves the question of why it is so persistent, given that it has been on record since 1948. Part of the answer is uncomfortably practical: the sentences that carry it are socially useful. They open conversations, insult nobody and give the other person material to narrate themselves with — which is why cold readers, consultants and fortune tellers keep reinventing them independently of one another. The other part is more fundamental: personal validation is not stupidity but the everyday form of confirmation bias, and confirmation bias belongs to the basic equipment of judgement. You do not get rid of it; you can only know it and, at the decisive points — in court, in diagnostics, in personnel selection — replace it with procedures that disarm it: blinding, comparison groups, matched rather than presented descriptions.
Forer’s original study is also one of the few in this archive that was practically never in doubt. The finding has been repeated for decades in dozens of variants — with students and with working adults, with horoscopes and with mock computer print-outs — and the mean assent lands reliably in the upper range of the scale. It replicates so well because it measures nothing exotic, but the normal operation of the human self-image.
What remains
Secure is the effect itself: general, benevolent, two-sided sentences are accepted as an apt self-description when they arrive as a personal evaluation from a credible source. Equally secure is what follows from it — that assent is no evidence of accuracy. What remains open, and interesting, is the reverse direction: how much of the assent to good assessment also rests on Barnum components has rarely been cleanly separated out. A test finding can be valid and at the same time be accepted for reasons that have nothing to do with its validity.
In everyday life a Barnum statement can be spotted with a single question: who could this sentence not apply to? If no answer comes, it describes nobody. In practice that means checking evaluations in the plural rather than the singular — swap sheets with your neighbour, cover the name, lay three interpretations side by side. In his subtitle Forer himself declared his paper a classroom demonstration of gullibility, not a finding about the particular condition of his students. The effect does not shame, it demonstrates; that is why the original is still repeated in teaching today, as one of the few experiments that are their own debriefing.
Sources. Forer, B. R. (1949). The fallacy of personal validation: A classroom demonstration of gullibility. Journal of Abnormal and Social Psychology, 44(1), 118–123. · Meehl, P. E. (1956). Wanted — a good cookbook. American Psychologist, 11, 263–272. · Snyder, C. R., Shenkel, R. J., & Lowery, C. R. (1977). Acceptance of personality interpretations: The “Barnum effect” and beyond. Journal of Consulting and Clinical Psychology, 45(1), 104–114. · Dickson, D. H., & Kelly, I. W. (1985). The 'Barnum effect' in personality assessment: A review of the literature. Psychological Reports, 57, 367–382. · Furnham, A., & Schofield, S. (1987). Accepting personality test feedback: A review of the Barnum effect. Current Psychology, 6(2), 162–178. · Carlson, S. (1985). A double-blind test of astrology. Nature, 318(6045), 419–425. · On the retest reliability of type tests, see the evidence sheet of this archive’s Jung file.
Sources, and why they are here
Forer, B. R. (1949). The fallacy of personal validation: A classroom demonstration of gullibility. Journal of Abnormal and Social Psychology, 44(1), 118–123.
The original study: 39 students, thirteen sentences, tables 1 and 2 of the data block — and the source of the quotation.
Meehl, P. E. (1956). Wanted — a good cookbook. American Psychologist, 11, 263–272.
Where the effect got its name, and it got it as a warning sign inside a plea for statistical rather than intuitive evaluation.
Snyder, C. R., Shenkel, R. J., & Lowery, C. R. (1977). Acceptance of personality interpretations: The „Barnum effect“ and beyond. Journal of Consulting and Clinical Psychology, 45(1), 104–114.
Takes the assent apart into its conditions — personalisation, authority, flattery — and thereby supplies the construction manual of the first section.
Dickson, D. H., & Kelly, I. W. (1985). The „Barnum effect“ in personality assessment: A review of the literature. Psychological Reports, 57, 367–382.
The first of the two reviews behind the finding that flattery is the most reliable of the four ingredients.
Furnham, A., & Schofield, S. (1987). Accepting personality test feedback: A review of the Barnum effect. Current Psychology, 6(2), 162–178.
The second review; it reaches the same result independently, including for unflattering descriptions.
Carlson, S. (1985). A double-blind test of astrology. Nature, 318(6045), 419–425.
The matching test applied to astrology, double-blind and agreed in advance with the professional associations — the procedure that separates assent from accuracy.
On the retest reliability of type tests, see the evidence sheet of this archive's Jung file.
That a substantial share of test-takers are assigned a different type on a second run is documented there and is not carried twice.