Case study
The profile that started this
Robby, August 2026
I took the IPIP-NEO-120 expecting it to tell me something I did not know. Instead it told me something I knew was wrong, and I could see exactly which question had made it wrong while I was answering it.
The instrument is good. It is public domain, it has been administered to hundreds of thousands of people, and its facet structure has held up for twenty-five years. Nothing that follows is a complaint about its quality. It is a complaint about a modelling assumption so deeply built in that no score can express its failure — and my profile happens to break it in three different places at once.
1. A domain score that describes nobody
My Neuroticism came out at the 7th percentile. Very low. Calm, even-tempered, hard to rattle — and for five of the six facets underneath it, that is accurate. Anxiety 12. Depression 6. Vulnerability 9. Anger 11.
Self-Consciousness: 94th.
That is an eighty-seven point spread inside one domain. The domain score is the mean of its six facets, and the mean of that set is a number that describes no part of me. It says I am not easily disturbed, which is true, and it buries the one way in which I very much am.
Johnson — who built this inventory — says plainly that when facets within a domain disagree, you should trust the facets. He is right, and almost no test that uses his items says so where anyone will read it, because saying it undercuts the five big numbers at the top of the page.
2. A score built by a habit of speech
Modesty: 97th percentile. I noticed this one while answering, because I could feel myself doing it. Every item asking me to claim something about myself got a slightly more deflating answer than the truth deserved. Not dishonesty — a reflex.
That reflex does not stay in the Modesty facet. It is applied to all 120 items, so every self-descriptive score on the whole instrument is pulled down a little by a response habit that the instrument then reports as a trait. The 97 is real. It is also a measurement artifact contaminating everything around it.
3. The one that became the whole project
Altruism: 33rd percentile. Sympathy: 38th.
I have spent a substantial part of my adult life caring for people. The score is not describing that, and my first instinct was that the test was simply wrong. It is not wrong. I went back and read the items.
I am concerned about others.
I sympathise with the homeless.
I feel sympathy for those who are worse off than myself.
Every one of them is written about people in general. And my honest answer to all of them is no.
My honest answer about the six people inside my circle is that I would do essentially anything. The instrument has no way to hold both of those answers, so it asked for one, and I gave the one the question was actually asking for. Then it averaged a distribution I do not have and reported the middle of it.
The other Agreeableness facets say what is really going on: Cooperation 71, Dutifulness 87, Morality 90, Modesty 97. That is a profile organised around duty and code rather than around vicarious feeling. The prosocial behaviour is entirely there. It is conditional and it runs on obligation rather than on empathy, and a single number for “Altruism” cannot say that.
The differential is not measurement error. It is the most descriptively important thing about that trait in me, and it is exactly what the instrument discards.
4. And one that is just an artifact
Openness: 25th percentile, which reads as incurious. Five of its six facets sit between 30 and 54 — unremarkable. The sixth, Emotionality, is at 3, and it drags the whole domain down on its own.
Openness/Emotionality measures awareness of and interest in one’s own feelings. It is a real facet and mine is genuinely low. It is also the least related to what most people hear in the word “openness”, and it is doing most of the work in that number.
The scores
Facet percentiles from a single IPIP-NEO-120 administration, August 2026. No item-level responses were retained, so this is a static record rather than a live result.
| Neuroticism | 7 | |
|---|---|---|
| Extraversion | 21 | |
| Openness | 25 | |
| Agreeableness | 60 | |
| Conscientiousness | 82 |
The facets that matter to the argument
| Anxiety | 12 | ||
|---|---|---|---|
| Anger | 11 | ||
| Depression | 6 | ||
| Self-Consciousness | 94 | the contradiction | |
| Immoderation | 22 | ||
| Vulnerability | 9 | ||
| Trust | 44 | ||
| Morality | 90 | ||
| Altruism | 33 | the item problem | |
| Cooperation | 71 | ||
| Modesty | 97 | response style | |
| Sympathy | 38 | the item problem | |
| Imagination | 41 | ||
| Artistic Interests | 30 | ||
| Emotionality | 3 | drags the domain | |
| Adventurousness | 54 | ||
| Intellect | 48 | ||
| Liberalism | 37 |
Percentiles are against Johnson’s 619,150-case reference sample, which is a self-selected internet sample rather than a probability sample. Good for structure and rough placement, inadequate for anything clinical. That caveat applies to every percentile on this site and is not specific to this profile.
What this profile is not evidence of
One person noticing something about their own test is an anecdote, and the fact that the anecdote is mine makes it worse, not better — I designed the instrument that was built to confirm it. That is exactly the shape of a finding that turns out to be nothing.
So the construct has pre-registered falsification conditions, the rival explanations are administered alongside it so they can actually win, and every gradient figure on this site is labelled provisional until a test-retest study says otherwise. If it fails, that result gets published here with the same prominence as this page.