The research rationale
One number cannot describe a step function
Standard personality instruments assume trait expression is roughly consistent across social contexts. For a meaningful minority of people that assumption is false — and there is no mechanism in the instrument to notice.
1The problem
Every widely used Big Five instrument reports a facet as a single number. That number is a claim: it asserts you have a characteristic level of the trait, and that variation around it is either small or uninteresting.
For most facets and most people, that claim is defensible. This project concerns the cases where it is not — and where the deviation is not noise but the most descriptively important feature of the trait in that person.
The clearest case is Agreeableness. Its Altruism and Sympathy items are written about people in general — the needy, strangers, humanity at large:
“I am concerned about others.”
“I sympathize with the homeless.”
“I feel sympathy for those who are worse off than myself.”
Someone whose warmth is strongly conditional on closeness — guarded toward strangers, close to unconditional toward a small circle — has no accurate way to answer. Any single response overstates one and understates the other. The instrument returns approximately the mean of a distribution the respondent does not have.
2This is not measurement error
Two people can score identically on Altruism and be completely different:
- Person A — consistently moderate helpfulness toward everyone.
- Person B — near-zero toward strangers, near-total toward the inner circle.
These are not the same person. They behave differently in most situations that matter — in caregiving, in crisis, in how they form relationships, in where finite attention goes. A measurement model that cannot distinguish them is not merely imprecise; it is describing somebody who does not exist.
It is worse than an average, and the difference matters
An average at least summarises something. The standard items are written about people in general, and that phrasing does not blur Person B’s two behaviours together — it presupposes that a single general answer exists for them. It does not. So the number the instrument returns is not an imprecise summary of Person B’s range; it is an answer to a question they cannot truthfully answer, and it lands on a value they do not occupy toward anybody.
This changes what we do with it. For a trait where the gradient is real, the results page leads with the ring reading and demotes the conventional percentile to what it still legitimately is: the only figure on the page with a real reference sample behind it, and the only one comparable to every other Big Five test. It is a comparability number, not the description.
3What we measure instead
The gradient, not the point.
Selected facets are administered four times, at four ordered levels of closeness — your inner circle, friends, acquaintances, strangers — and the shape of the resulting curve is the object of measurement.
expression
^
7 | *----* flat — the usual model fits
|
5 | *----*
| \ tapers — ordinary modulation
3 | *----*
| *----*
1 | \____*----* step — a boundary, and it matters
+----------------------------> which side you are on
inner friends acq. strangersFour points rather than two, because a two-point difference score is statistically unusable — its reliability lands around .27 with our item counts, and can go negative. Four ordered points support an exact decomposition into level, slope and curvature whose precision can be calculated rather than assumed.
4The same failure, one level further down
Each measured trait is asked about three different ways, and those three do not always agree with each other. Someone whose embarrassment in front of strangers is extreme, while the other two statements about the same trait stay level, produces a facet curve that is the mean of a disagreement — a gentle slope that looks unremarkable, with the interesting part averaged away.
That is the identical error the instrument was built to catch, occurring inside a single facet rather than across a person. So the results page draws the spread of the three statements behind every curve, always, rather than only when a threshold is crossed. Where they agree the band collapses onto the line; where they pull apart it opens, and the eye lands on the disagreement without needing to be told.
Statements that point opposite ways are a different finding
5Which traits, and why those
Ten facets, chosen so the central hypothesis can fail. Three are predicted flat and one is predicted to run backwards — without those, a gradient would be indistinguishable from someone simply agreeing with everything.
| Trait | Prediction | Why it is in the set |
|---|---|---|
| Trust | steep | An affective stance rather than a rule. Expected to gate sharply on closeness. |
| Altruism | steep | The facet that motivated the whole instrument: IPIP items ask about people in general. |
| Sympathy | steep | Vicarious feeling, which plausibly requires a relationship to activate at all. |
| Friendliness | steep | Warmth toward a target is almost definitionally target-dependent. |
| Gregariousness | moderate | A crowd of strangers and a room of friends are plainly different experiences, and standard items rarely disambiguate. |
| Self-Consciousness | open | NO DIRECTIONAL PREDICTION. Whether shame is sharper before strangers or before people whose judgement you cannot escape is genuinely open, and either answer is informative. |
| Morality | flat | CONTROL. A conduct rule, applied uniformly. If this gates as much as Altruism, H1 is false. |
| Cooperation | flat | CONTROL. Conflict conduct rather than affect. |
| Dutifulness | flat | CONTROL. Obligation as a rule about oneself rather than a response to a person. |
| Anger | reversed | REVERSAL CONTROL, and the strongest available defence against response-style artifacts. Predicted to be expressed MORE toward close others -- you are short with family and civil with strangers -- consistent with Clifton (2014). A uniform response bias cannot produce opposite-signed slopes. |
The reversal is the important one
6What is already known
This section exists because the honest answer to “is this new?” is partly, and the parts that are not new should be said first.
Clifton (2014) had people rate their Big Five as experienced with each of thirty individual members of their social network, and found that expression varied with how central the other person was. A closeness gradient in Big Five expression has therefore already been demonstrated. No claim of novelty for the phenomenon is available here, and none is made.
Repeating one item set across several relational targets is also established — Fraley et al.’s ECR-RS does it with four attachment targets, and self-concept differentiation (Donahue et al., 1993) has measured cross-role variability since the nineties.
What is plausibly new is narrower:
- Facet-level resolution. Prior work is domain-level or adjective-level.
- Signed, directional profiles. Self-concept differentiation discards direction by construction — and direction is exactly the information we want.
- Closeness as the ordering axis rather than social role, which connects trait expression to the discounting and moral-circle literatures rather than to role theory.
7The finding that could kill this
Lenhausen, Bleidorn & Hopwood (2023), Assessment
Taken at face value, that result predicts this instrument will find nothing. We cite it ourselves, prominently, because a reviewer who finds it first will use it to end the conversation.
The distinction we depend on is this: they manipulated the comparison standard — rate yourself relative to close others. We manipulate the behavioural target — how you actually behave toward close others. Those are different psychological operations. The first asks you to re-anchor a self-assessment; the second asks about a different set of behaviours.
Frame-of-reference research that manipulates target or domain does move scores substantially — a meta-analysis found contextualised measures achieve criterion validity of .24 against .11 for non-contextualised ones.
But the distinction is carried entirely by item wording, which makes it an engineering constraint rather than an argument. Every stem in our instrument is written as behaviour toward a target and must be unreadable as a comparison. You can read all of them on the methods page.
If our gradients also come out at d < .05, the honest conclusion is that Lenhausen et al. were measuring the same thing, and the construct does not exist. We will say so.
8What would prove us wrong
Recorded before any data was collected, so they cannot be revised afterwards. The construct is not supported if any of these holds:
- The frames do not separate. Ring-framed scores differ by d < .05, replicating the null above.
- No differentiation across traits. The conduct controls gate as much as Altruism, and Anger does not reverse. That would indicate a response-style artifact rather than gating.
- Measurement non-invariance across rings. Without it, a between-ring difference reflects the items behaving differently rather than the person differing.
- Inadequate reliability. If the combined index is not stable enough for individual interpretation, we withdraw individual feedback and report population findings only.
- No incremental validity. If contraction predicts nothing beyond overall trait level, self-concept differentiation, neuroticism, and existing closeness measures, then it reduces to things already measurable at a fraction of the item cost.
9What this is not
- Not IRB-approved institutional research. An independent project, run openly.
- Not clinical, not diagnostic, not a screen for anything.
- Never for hiring, admissions, or any evaluative decision about another person.
- Not population norms. The reference sample is self-selected internet volunteers — young, English-speaking, and not representative. Every percentile on this site says which group it refers to.
There is also a well-documented trap in reports like these: people accept randomly generated personality descriptions as accurate about themselves, and the vaguer and more flattering the prose, the higher the acceptance. That is why our interpretation stays tied to specific numbers, hedged by their uncertainty, and occasionally unflattering.