Why self-assessments lie to you
The result felt accurate. That is the problem, and it has been the problem since 1949.
In 1949 Bertram Forer gave 39 students a personality test, then handed each one a page describing their individual results. He asked them to rate how accurate the description was. The average came back at 4.30 out of 5.
Every student had been given the identical page. Forer had assembled it by cutting statements out of a newsstand astrology book.
What that experiment actually establishes
Not that his students were foolish. The result has been repeated for decades and the average sits near the same place. Michel Gauquelin later mailed thousands of readers of a French magazine one identical horoscope, and most of them called it accurate. It was the birth chart of Marcel Petiot, a doctor convicted of murder.
The finding is about the instrument, not the reader. It says that a person's agreement with a description of themselves carries almost no information about whether the description is true.
Which means the sentence this is so me, the one every personality product is designed to produce, is not evidence of accuracy. It is evidence that the writing was agreeable, and agreeable is easy.
A person's agreement with a description of themselves carries almost no information about whether the description is true.
How the trick works, in sentences
Forer's page worked because of how the statements were built. Look at the shapes, because once you can see them you will find them in every test you have taken.
The statement with no opposite
Try to picture the person for whom this is false. You cannot, so agreeing costs you nothing.
SayYou have a great need for other people to like and admire you.
The statement with both poles
This is the highest-scoring shape on record, and the reason is structural. There is no observation that could contradict it.
SayAt times you are extroverted and sociable, while at other times you are reserved.
The claim about your interior
Nothing can disconfirm this, including your denial, which the sentence has already framed as the outside self talking.
SayDisciplined outside, you tend to be insecure inside.
The flattering claim
Favourable statements get accepted more readily than unfavourable ones at the same level of vagueness. So a result that reads as accurate because it is generous has told you nothing.
The expensive versions do it too
This is not only a problem with free tests. A widely sold strengths report contains the sentence: perhaps you enjoy helping individuals understand complicated processes, contracts, or regulations they might not grasp without your assistance.
That sentence cannot be wrong. The hedge makes it optional, the list of three nouns widens the net until something lands, and the reader supplies the match. It sits inside a respected, well-researched, expensive instrument.
Nobody drifts into this voice through laziness. Writers drift into it because it is safe, universally agreeable, and never generates a complaint.
01
What you say about yourself
one source. Every other test in this category stops here.
02
What five people say
answered separately, without seeing each other or you
03
The gap between them
the only thing here that is not available from the inside
The deeper problem is the source, not the wording
Suppose a test avoided every trick above. Suppose every sentence was sharp, falsifiable and specific. It would still have one source of information about you, and that source would be you.
Kruger and Dunning established in 1999 that people least skilled at something are the worst judges of their own skill at it. Assessing the work requires the ability they lack. The blind spot and the thing hiding in it are the same size by definition.
There is a second problem that has nothing to do with competence. You experience your intentions and other people experience your behaviour. You know you meant the short reply kindly. Nobody else was given that information, and the reply is all they got.
Connelly and Ones, in a 2010 meta-analysis in Psychological Bulletin, examined ratings made by people who know the person being rated. Those ratings carry information that self-ratings do not, and for some outcomes they predict better than a person's own answers.
The effect grows with the number of raters and with how well they know you. That is the entire finding, and it is enough: the useful information about how you come across is distributed across several people, and none of them is you.
The mechanism, and it is not dishonesty
Nobody rating themselves is trying to cheat. The comparison is rigged before the question is asked, because of what each side has access to.
You rate your own patience against every time you wanted to snap and did not. Those near misses are the bulk of your evidence and they are invisible to everyone else. Your colleague rates your patience against the two occasions you did snap, because those are the only ones they were in the room for.
So you are scoring your intentions and they are scoring your behaviour. Both answers are honest. They are answers to different questions, and only one of them is the question you actually wanted answered.
This is why arguing with a low rating never resolves anything. You reach for the restraint you exercised, which is real, and they cannot check it, and neither of you is lying.
What a falsifiable sentence about a person looks like
The repair is not more honesty. It is a sentence somebody could be wrong about. Three rewrites, each moving from a claim with no complement to one with a witness.
Vague. You value clear communication. Sharp. You send the follow-up message that repeats what was agreed, and the people who were already clear read it as being checked up on.
Vague. You are highly driven. Sharp. You have moved a deadline forward twice this year without being asked, and the people working to it found out after it changed.
Vague. You care deeply about your close relationships. Sharp. You go three weeks without contacting anyone, and you expect the friendship to be where you left it.
Run the same check on any result you have been given. Name one person for whom the sentence is false. If no such person exists, the sentence is about everybody, and a sentence about everybody told you nothing about you.
What the tests are legitimately good for
They are not worthless, and pretending otherwise would be its own kind of dishonesty. They give a group of people a shared vocabulary for differences that were previously just friction.
A team that can say I need the agenda in advance, without it sounding like a complaint, is a team that has gained something real. The label did the work of making a preference discussable. That is a genuine use.
They are also a decent conversation starter, which is not a small thing. Reading a description of yourself out loud to somebody who knows you tends to produce the sentence that matters, which is usually them saying no, not that bit.
What none of that establishes is accuracy. A vocabulary can be useful and wrong at the same time, and a conversation starter is judged by the conversation, not by the starter.
The part that costs us to say
Everything above applies to us. If we write a result and you tell us it is accurate, we have learned nothing, because Forer's students said the same thing about astrology filler.
So agreement is not the test we intend to be judged on. The real test is whether somebody can pick their own result out of a lineup of plausible wrong ones. If people cannot do that at well above chance, our writing is Barnum prose regardless of how accurate it felt, and the voice has failed.
We think that is the right test and we intend to run it. It can come out against us. The rest of what we refuse to claim is written in the same spirit.
Keep taking tests if you enjoy them. They are a decent vocabulary for talking about yourself and a bad instrument for finding out about yourself.
When you read a result, run one check on each sentence: can you describe a real person for whom this is false? If you cannot, the sentence is not about you. It is about everybody, which is another way of saying it is about nobody.
Forer, B. R. (1949). The fallacy of personal validation. Journal of Abnormal and Social Psychology, 44(1), 118 to 123. Mean accuracy 4.30 out of 5, n=39. Kruger, J. and Dunning, D. (1999). Journal of Personality and Social Psychology, 77(6). Connelly, B. S. and Ones, D. S. (2010). An other perspective on personality. Psychological Bulletin, 136(6).
The part this page cannot do
This page can tell you that your own answers are the wrong instrument. It cannot tell you what the right instrument would have said.
The information you are looking for exists, and it is distributed across a handful of people who have never been asked in a way that made answering safe.
Start the interviewIt takes about fifteen minutes and it is free. The $29 Portrait is only charged after five people finish and the evidence clears.