Why self-assessments lie to you

The result felt accurate. That is the problem, and it has been the problem since 1949.

In 1949 Bertram Forer gave 39 students a personality test, then handed each one a page describing their individual results. He asked them to rate how accurate the description was. The average came back at 4.30 out of 5.

Every student had been given the identical page. Forer had assembled it by cutting statements out of a newsstand astrology book.

What that experiment actually establishes

Not that his students were foolish. The result has been repeated for decades and the average sits near the same place. Michel Gauquelin later mailed thousands of readers of a French magazine one identical horoscope, and most of them called it accurate. It was the birth chart of Marcel Petiot, a doctor convicted of murder.

The finding is about the instrument, not the reader. It says that a person's agreement with a description of themselves carries almost no information about whether the description is true.

Which means the sentence this is so me, the one every personality product is designed to produce, is not evidence of accuracy. It is evidence that the writing was agreeable, and agreeable is easy.

A person's agreement with a description of themselves carries almost no information about whether the description is true.

How the trick works, in sentences

Forer's page worked because of how the statements were built. Look at the shapes, because once you can see them you will find them in every test you have taken.

  1. The statement with no opposite

    Try to picture the person for whom this is false. You cannot, so agreeing costs you nothing.

    SayYou have a great need for other people to like and admire you.

  2. The statement with both poles

    This is the highest-scoring shape on record, and the reason is structural. There is no observation that could contradict it.

    SayAt times you are extroverted and sociable, while at other times you are reserved.

  3. The claim about your interior

    Nothing can disconfirm this, including your denial, which the sentence has already framed as the outside self talking.

    SayDisciplined outside, you tend to be insecure inside.

  4. The flattering claim

    Favourable statements get accepted more readily than unfavourable ones at the same level of vagueness. So a result that reads as accurate because it is generous has told you nothing.

The expensive versions do it too

This is not only a problem with free tests. A widely sold strengths report contains the sentence: perhaps you enjoy helping individuals understand complicated processes, contracts, or regulations they might not grasp without your assistance.

That sentence cannot be wrong. The hedge makes it optional, the list of three nouns widens the net until something lands, and the reader supplies the match. It sits inside a respected, well-researched, expensive instrument.

Nobody drifts into this voice through laziness. Writers drift into it because it is safe, universally agreeable, and never generates a complaint.

One self-report against five independent reportsOn the left, a single circle labelled what you say about yourself, which is the only source every other test in this category has. On the right, five separate circles, one for each person who answers about you without seeing the others. The space between them is labelled the gap, which is what the Portrait is written about.Youone sourceevery other test stops here12345five people, answering separatelyTHE GAP
  1. 01

    What you say about yourself

    one source. Every other test in this category stops here.

  2. 02

    What five people say

    answered separately, without seeing each other or you

  3. 03

    The gap between them

    the only thing here that is not available from the inside

Every other test in this category is the single circle: one source, which is you. The five never see each other's answers, and the product is the gap between the one and the five.

The deeper problem is the source, not the wording

Suppose a test avoided every trick above. Suppose every sentence was sharp, falsifiable and specific. It would still have one source of information about you, and that source would be you.

Kruger and Dunning established in 1999 that people least skilled at something are the worst judges of their own skill at it. Assessing the work requires the ability they lack. The blind spot and the thing hiding in it are the same size by definition.

There is a second problem that has nothing to do with competence. You experience your intentions and other people experience your behaviour. You know you meant the short reply kindly. Nobody else was given that information, and the reply is all they got.

Connelly and Ones, in a 2010 meta-analysis in Psychological Bulletin, examined ratings made by people who know the person being rated. Those ratings carry information that self-ratings do not, and for some outcomes they predict better than a person's own answers.

The effect grows with the number of raters and with how well they know you. That is the entire finding, and it is enough: the useful information about how you come across is distributed across several people, and none of them is you.

The mechanism, and it is not dishonesty

Nobody rating themselves is trying to cheat. The comparison is rigged before the question is asked, because of what each side has access to.

You rate your own patience against every time you wanted to snap and did not. Those near misses are the bulk of your evidence and they are invisible to everyone else. Your colleague rates your patience against the two occasions you did snap, because those are the only ones they were in the room for.

So you are scoring your intentions and they are scoring your behaviour. Both answers are honest. They are answers to different questions, and only one of them is the question you actually wanted answered.

This is why arguing with a low rating never resolves anything. You reach for the restraint you exercised, which is real, and they cannot check it, and neither of you is lying.

What a falsifiable sentence about a person looks like

The repair is not more honesty. It is a sentence somebody could be wrong about. Three rewrites, each moving from a claim with no complement to one with a witness.

Vague. You value clear communication. Sharp. You send the follow-up message that repeats what was agreed, and the people who were already clear read it as being checked up on.

Vague. You are highly driven. Sharp. You have moved a deadline forward twice this year without being asked, and the people working to it found out after it changed.

Vague. You care deeply about your close relationships. Sharp. You go three weeks without contacting anyone, and you expect the friendship to be where you left it.

Run the same check on any result you have been given. Name one person for whom the sentence is false. If no such person exists, the sentence is about everybody, and a sentence about everybody told you nothing about you.

What the tests are legitimately good for

They are not worthless, and pretending otherwise would be its own kind of dishonesty. They give a group of people a shared vocabulary for differences that were previously just friction.

A team that can say I need the agenda in advance, without it sounding like a complaint, is a team that has gained something real. The label did the work of making a preference discussable. That is a genuine use.

They are also a decent conversation starter, which is not a small thing. Reading a description of yourself out loud to somebody who knows you tends to produce the sentence that matters, which is usually them saying no, not that bit.

What none of that establishes is accuracy. A vocabulary can be useful and wrong at the same time, and a conversation starter is judged by the conversation, not by the starter.

The part that costs us to say

Everything above applies to us. If we write a result and you tell us it is accurate, we have learned nothing, because Forer's students said the same thing about astrology filler.

So agreement is not the test we intend to be judged on. The real test is whether somebody can pick their own result out of a lineup of plausible wrong ones. If people cannot do that at well above chance, our writing is Barnum prose regardless of how accurate it felt, and the voice has failed.

We think that is the right test and we intend to run it. It can come out against us. The rest of what we refuse to claim is written in the same spirit.

Keep taking tests if you enjoy them. They are a decent vocabulary for talking about yourself and a bad instrument for finding out about yourself.

When you read a result, run one check on each sentence: can you describe a real person for whom this is false? If you cannot, the sentence is not about you. It is about everybody, which is another way of saying it is about nobody.

Forer, B. R. (1949). The fallacy of personal validation. Journal of Abnormal and Social Psychology, 44(1), 118 to 123. Mean accuracy 4.30 out of 5, n=39. Kruger, J. and Dunning, D. (1999). Journal of Personality and Social Psychology, 77(6). Connelly, B. S. and Ones, D. S. (2010). An other perspective on personality. Psychological Bulletin, 136(6).

The part this page cannot do

This page can tell you that your own answers are the wrong instrument. It cannot tell you what the right instrument would have said.

The information you are looking for exists, and it is distributed across a handful of people who have never been asked in a way that made answering safe.

Start the interview

It takes about fifteen minutes and it is free. The $29 Portrait is only charged after five people finish and the evidence clears.