Skip to content
AnamneseArchive of Psychology

At least two characters

The Journal29 August 20267 min read

The face was never a piece of evidence

From physiognomy to the AI that renews the same promise

The same grid, twice — and no face beneath it. On the right stands what is actually in dispute: for the 2016 paper the provenance of the image collections, for the 2018 paper not the figures but what was measured. Agreement still is not accuracy.Drawing by the archive

In November 2016 Xiaolin Wu and Xi Zhang posted a paper online claiming that criminality could be read from portrait photographs. There were 1,856 faces, just under half of them convicted offenders, and four classifiers, all with good results. The authors even named the features: the curvature of the lip, the distance between the inner corners of the eyes, the angle between nose and mouth. And they reported that the faces of non-criminals resembled one another more closely — a “law of normality” for faces.

The promise is old, the computing power new. Where physiognomy held a brass ruler to the forehead, today a neural network sits between photograph and judgment. The leap of logic has stayed the same.

What forms in a tenth of a second

The starting finding is solid, and it is what makes the matter interesting. People form judgments about faces in roughly 100 milliseconds; those allowed to look longer mainly become more confident, not different. The impressions are therefore not laboriously worked out but reflexive — they feel like perception rather than interpretation.

This is exactly why a single photograph misleads so reliably. Alexander Todorov and Jenny M. Porter showed that pictures of the same person can produce impressions as different as pictures of different people. The variation within one person matches the variation between people, or exceeds it. Lighting, expression, angle and the choice of image alter apparent character — so the choice of photograph alters the result of any analysis that rests on a photograph.

Agreement is not accuracy

Many observers find the same face competent or trustworthy. That proves they share similar visual rules — not that the person possesses the trait. A faint angry expression can be overgeneralised as dominance, a round face as harmlessness. Cultural learning makes the judgment consistent and potentially unfair.

How far consistency stands from accuracy has been measured by a meta-analysis of perceived trustworthiness. At the level of individual faces the relation between impression and actual behaviour came to r = .14, at the level of individual observers to r = .27. The authors regard this as explicitly unusable in practice. It is enough to see a signal in large samples, and far too little to decide about a person.

Two algorithms and what they actually measured

Wu and Zhang’s criminality classifier took the path such work usually takes: first attention, then contradiction. Kevin W. Bowyer, Michael King, Walter Scheirer and Kushal Vangara consider the whole approach hopeless and the promising results an illusion arising from an inadequate experimental design. That is precisely where the dispute lies: Wu and Zhang state that they controlled their images for origin, sex, age and facial expression. The critics regard this control as insufficient — and because both sides accept the same hit rates, the matter is settled not by statistics but by the question of where the two collections of images came from. Wu and Zhang themselves added a postscript in which they described their work as purely academic discussion and expressed surprise at its public reception.

The second case weighs more heavily because it appeared in a journal. In 2018, in the Journal of Personality and Social Psychology, Yilun Wang and Michal Kosinski analysed 35,326 facial images. Their classifier, they reported, correctly distinguished gay from heterosexual men in 81% of cases, and women in 71%. With five images per person the figures rose to 91% and 83%; human judges reached 61% and 54%.

The decisive detail stands in the paper itself. Among the features the classifier used, the authors list, alongside fixed traits such as the shape of the nose, explicitly variable ones: expression and grooming style — that is, hairstyle, facial hair, glasses, cosmetics. These are not properties of a face but decisions by a person about how they wish to appear in a self-uploaded portrait. What was measured was self-presentation in a particular social setting, not physiognomy.

The authors themselves read their finding differently. Gay men and lesbian women, they write, tend to have gender-atypical facial morphology, expression and grooming styles — in line with the theory of prenatal hormone effects. What they report in the same breath is remarkable: models aimed at gender atypicality alone identified gay men in 57% and lesbian women in 58% of cases. By far the largest part of the advantage therefore comes from something other than mere gender atypicality — and the paper itself names grooming style and expression as part of that something.

Andrew Gelman, Greggor Mattson and Daniel Simpson have given this a name: the fallacy of decontextualised measurement.

the ideals of scientific precision strip the context from intrinsically social phenomena

Andrew Gelman, Greggor Mattson & Daniel Simpson, Gaydar and the Fallacy of Decontextualized Measurement (2018)

The assumption that there is a purely biological measure for the perception of sexual orientation lifts a social phenomenon out of its context and sells the precision of the statistics as the precision of the claim. Anyone who reads the figure 91% thinks of a property. What was measured was a data set.

When prediction does not mean truth

An algorithm can score hits in one data set and still learn the wrong feature. It may be detecting lighting, camera type, clothing or social inequality instead of personality. If the model does not work in another population, its apparent insight into people was only a shortcut through the training data.

Even a stable prediction proves no innate property of the face. Discriminatory treatment can link appearance, expression and documented life courses: whoever is stopped more often turns up more often in records, and whoever turns up in records becomes a training example. An algorithm that learns past hiring, lending or sentencing decisions reliably predicts how decisions used to be made.

This becomes especially dangerous in policing, hiring and insurance. An untestable judgment of character acquires the appearance of objective precision while those affected know neither the cause nor the remedy. What is needed is therefore external validation, error rates reported separately for different groups, and the question of whether such a classification may legitimately be deployed at all.

TestYearScopeWhat came of it
Willis & Todorov2006five experiments100 ms suffice for a judgement from a face
Todorov & Porter2014pictures of the same personvariation within a person matches that between persons
Wu & Zhang, criminality20161,856 portraitshit rates undisputed — the provenance of the images not
Wang & Kosinski201835,326 facial images81 % and 71 %; on gender atypicality alone only 57 % and 58 %
Bowyer et al., re-examination2020critique of the designthe results an artefact, not a finding
Foo et al., meta-analysis2022trustworthinessr = .14 per face, r = .27 per observer

What remains

Faces are not meaningless. They show momentary expression, attention and some states, and these signals depend on context. What holds is the fast, shared, largely inaccurate impression: it forms in a tenth of a second, many people share it, and it says almost nothing about the person. What does not hold is the step from that impression to a trait — and no model changes this, because the problem does not lie in the computing power but in the inference.

What remains contested is how to deal with models that do find a signal. That a classifier works in one data set is an empirical statement about that data set; whether a property or a context lies behind it is decided only by external testing. Both sides of the debate around Wang and Kosinski appeal to the same numbers.

In everyday life the fallacy shows itself in a missing sentence. Where someone reads a face and does not say how the judgment could be checked, it is not a judgment but an impression with a job title. Good analysis of people observes behaviour across situations and time, uses self- and informant reports, and asks about concrete actions. Physiognomy does not fail because faces mean nothing. It fails because it turns meaning into destiny.

Sources, and why they are here

  1. Willis, J., & Todorov, A. (2006). First impressions: Making up your mind after a 100-ms exposure to a face. Psychological Science, 17(7), 592–598.

    The solid starting finding, without which the article would have no tension: the judgement is there before the thinking begins — and more time only makes it more confident, not different.

  2. Todorov, A., & Porter, J. M. (2014). Misleading first impressions: Different for different facial images of the same person. Psychological Science, 25(7), 1404–1417.

    The finding that strikes every single-photograph analysis at its core: pictures of the same person vary as widely as pictures of different people. Whoever chooses the photograph chooses part of the result.

  3. Todorov, A., Olivola, C. Y., Dotsch, R., & Mende-Siedlecki, P. (2015). Social attributions from faces: Determinants, consequences, accuracy, and functional significance. Annual Review of Psychology, 66, 519–545.

    The review the overgeneralisations come from — the slightly angry expression read as dominance, the round face as harmlessness.

  4. Foo, Y. Z., Sutherland, C. A. M., Burton, N. S., Nakagawa, S., & Rhodes, G. (2022). Accuracy in facial trustworthiness impressions: Kernel of truth or modern physiognomy? A meta-analysis. Personality and Social Psychology Bulletin, 48(11), 1580–1596.

    Puts a number on the gap between agreement and accuracy: r = .14 per face, r = .27 per observer — expressly declared by the authors to be of no practical use.

  5. Wu, X., & Zhang, X. (2016). Automated inference on criminality using face images. arXiv:1611.04135.

    The first of the two cases. What is remarkable is not the hit rate but that the authors name the very measures the nineteenth century already took.

  6. Bowyer, K. W., King, M. C., Scheirer, W. J., & Vangara, K. (2020). The „criminality from face“ illusion. IEEE Transactions on Technology and Society, 1(4), 175–183.

    The rejoinder that puts the quarrel in the right place: what is disputed is not the statistics but the design — that is, the provenance of the two image collections.

  7. Wang, Y., & Kosinski, M. (2018). Deep neural networks are more accurate than humans at detecting sexual orientation from facial images. Journal of Personality and Social Psychology, 114(2), 246–257.

    The weightier case, because it passed peer review. The decisive detail stands in the paper itself: the features used include grooming style and expression — decisions, not facial features.

  8. Gelman, A., Mattson, G., & Simpson, D. (2018). Gaydar and the fallacy of decontextualized measurement. Sociological Science, 5, 270–280.

    Gives the error its name, and the quotation comes from its abstract: the ideals of scientific precision strip a social phenomenon of its context.