The population
Explore the data
Everything an individual result compares you against, shown directly. Same items, same scoring engine, read across everyone rather than one person. Each figure carries the number of people behind it and the date it describes, because a statistic without either is a decoration.
Showing everyone. You can combine up to 3 filters — beyond that, a group can be small enough to identify someone even when the count looks large.
How Neuroticism is distributed
The steps are real. A facet score is the sum of four whole-numbered answers, so only certain values are possible and a smooth curve would be inventing shape the data does not have. Height is the share of people at each score, not a count, so a filtered group stays comparable with the whole population behind it.
Neuroticism across groups
Averages with their intervals. Sex and age differences in personality are real and small — usually a fraction of the spread within any one group — so the intervals are the point, not the dots.
Each bar is a 95% interval around that group’s average. Where two intervals overlap, the difference between those groups is not established by this data — however far apart the dots look. The axis is zoomed to the range the averages occupy, which is a small slice of the full 4–20 scale.
All thirty narrow traits
Every facet on one screen, drawn on a shared scale so the shapes can be compared directly rather than one at a time from memory. Select any of them to read what it measures.
Neuroticism
Extraversion
Openness
Agreeableness
Neuroticism across the age bands
The whole distribution for each age group, not just its average. Two groups can share a mean and be shaped completely differently — one clustered tightly, one split in two — and a table of averages hides exactly that. Every curve is drawn on the same axis and the same vertical scale, so both position and spread can be read across rows.
Show the numbers
| age band | n | mean | SD | α |
|---|---|---|---|---|
| 18-19 | 107,072 | 11.49 | 2.51 | 0.89 |
| 20-21 | 100,876 | 11.17 | 2.56 | 0.90 |
| 22-25 | 99,228 | 11.27 | 2.66 | 0.90 |
| 26-30 | 68,088 | 11.27 | 2.73 | 0.91 |
| 31-39 | 68,451 | 11.12 | 2.74 | 0.91 |
| 40+ | 68,570 | 10.55 | 2.66 | 0.90 |
The five domains by age
Where each domain sits in each age band, with the interval around every average. The intervals are the point: when they overlap, the difference between two bands is not something this sample can distinguish, however suggestive the line between them looks.
Show the numbers
| Trait | 18-19 | 20-21 | 22-25 | 26-30 | 31-39 | 40+ |
|---|---|---|---|---|---|---|
| Neuroticism | 11.49 ±0.02 | 11.17 ±0.02 | 11.27 ±0.02 | 11.27 ±0.02 | 11.12 ±0.02 | 10.55 ±0.02 |
| Extraversion | 13.87 ±0.01 | 13.93 ±0.01 | 13.70 ±0.01 | 13.50 ±0.02 | 13.39 ±0.02 | 13.27 ±0.02 |
| Openness | 13.67 ±0.01 | 13.63 ±0.01 | 14.00 ±0.01 | 14.11 ±0.02 | 13.96 ±0.02 | 13.91 ±0.02 |
| Agreeableness | 14.61 ±0.01 | 14.67 ±0.01 | 14.65 ±0.01 | 14.76 ±0.02 | 14.97 ±0.01 | 15.51 ±0.01 |
| Conscientiousness | 14.08 ±0.01 | 14.56 ±0.01 | 14.58 ±0.01 | 14.82 ±0.02 | 15.10 ±0.02 | 15.62 ±0.02 |
These are different people measured once, not the same people followed over time. A difference between two bands may be change that comes with age, or it may be that people born thirty years apart answer differently — a single sweep cannot tell those apart, and nothing on this page should be read as a claim that anyone changes.
Where women and men differ
Every facet, sorted by the size of the difference rather than by whether one exists. Differences are shown in standard deviations, because a gap of half a point means something different on a narrow trait than on a wide one and these thirty scales have to be comparable to each other.
Show the numbers
| Scale | d | 95% low | 95% high | n |
|---|---|---|---|---|
| Emotionality | 0.614 | 0.608 | 0.620 | 512,080 |
| Sympathy | 0.541 | 0.535 | 0.547 | 512,133 |
| Anxiety | 0.535 | 0.530 | 0.541 | 512,223 |
| Vulnerability | 0.521 | 0.515 | 0.527 | 512,097 |
| Altruism | 0.503 | 0.498 | 0.509 | 512,141 |
| Modesty | 0.450 | 0.444 | 0.456 | 512,113 |
| Morality | 0.385 | 0.380 | 0.391 | 512,075 |
| Artistic Interests | 0.313 | 0.307 | 0.318 | 512,113 |
| Cooperation | 0.290 | 0.284 | 0.295 | 512,074 |
| Achievement-Striving | 0.238 | 0.233 | 0.244 | 512,054 |
| Activity Level | 0.237 | 0.231 | 0.242 | 512,075 |
| Dutifulness | 0.194 | 0.189 | 0.200 | 512,051 |
| Anger | 0.189 | 0.183 | 0.194 | 512,162 |
| Depression | 0.175 | 0.169 | 0.181 | 512,101 |
| Immoderation | 0.160 | 0.155 | 0.166 | 512,044 |
| Cheerfulness | 0.131 | 0.125 | 0.136 | 512,012 |
| Liberalism | 0.120 | 0.115 | 0.126 | 511,998 |
| Orderliness | 0.098 | 0.093 | 0.104 | 512,124 |
| Self-Discipline | 0.097 | 0.092 | 0.103 | 512,080 |
| Self-Consciousness | 0.092 | 0.087 | 0.098 | 512,083 |
| Friendliness | 0.089 | 0.083 | 0.094 | 512,193 |
| Gregariousness | 0.074 | 0.069 | 0.080 | 512,124 |
| Trust | 0.020 | 0.014 | 0.026 | 512,185 |
| Self-Efficacy | -0.026 | -0.031 | -0.020 | 512,137 |
| Imagination | -0.061 | -0.067 | -0.055 | 512,178 |
| Cautiousness | -0.064 | -0.069 | -0.058 | 512,155 |
| Adventurousness | -0.113 | -0.119 | -0.107 | 512,103 |
| Assertiveness | -0.142 | -0.147 | -0.136 | 512,080 |
| Excitement-Seeking | -0.205 | -0.211 | -0.200 | 512,032 |
| Intellect | -0.230 | -0.236 | -0.225 | 512,046 |
With a large enough sample almost every difference becomes detectable, which is why the shaded strip matters more than the bars: inside it, a difference is real in the sense that it is not zero, and too small to tell you anything about a person in front of you. Group averages describe groups. The overlap between two of these distributions is always far larger than the gap between their centres.
How the traits relate
Every facet against every other facet. This is the structure the five broad domains are a summary of — and the places where the summary fails are visible here and nowhere else on the site.
The correlation matrix is computed over our own respondents, and no aggregate release has been published yet.
Nothing is hidden to make the page look tidier. Try removing a filter.
Where a single number is doing the most work
Everything above this point is the conventional Big Five: one number per trait, with an interval around it. That interval covers measurement error — how much the number would move if you answered the same statements again. It does not cover the other thing that moves it, which is who the statements are about.
The public-domain items behind several of these facets ask about people in general. “I go out of my way to help others.” A person whose answer is alwaysfor the four people closest to them and rarely for anyone else has no honest way to answer, and the instrument records the average of two behaviours rather than either one. The number is a reading taken at one unstated social distance and reported as though distance were irrelevant — so its interval is too narrow, and not by a rounding amount.
That error also does not wash out when six facets are averaged into a domain. Measurement error is roughly independent between facets, so it shrinks as they combine. Withholding is not independent: the whole premise is that it is a property of the person, so someone who reserves Altruism for their inner circle plausibly reserves Sympathy too. Correlated components survive averaging, which is how a facet-level problem becomes a domain-level one.
How much of each domain carries a closeness measurement
Coverage, not magnitude. A filled cell means closeness statements were written for that facet, so the question can be asked of it. Nobody has answered them yet, so a domain can be five-sixths covered and still turn out perfectly steady — which is what the controls below are for.
Agreeableness is the domain most exposed, at five facets of six — which is unsurprising, because it is the domain that motivated building any of this. Openness is not covered at all. And the interesting design detail is that Agreeableness is exposed in both directions: three of its five are predicted to gate steeply, and two are controls predicted not to move. If the controls gate as hard as the rest, the hypothesis is wrong, and it can be tested inside a single domain without comparing across domains at all.
The ten facets, and what was predicted for each
Neuroticism
- Angerpredicted: risesreversal control
REVERSAL CONTROL, and the strongest available defence against response-style artifacts. Predicted to be expressed MORE toward close others -- you are short with family and civil with strangers -- consistent with Clifton (2014). A uniform response bias cannot produce opposite-signed slopes. A control exists to be wrong. If this one falls with distance like the rest, the pattern is a response habit rather than a finding.
- Self-Consciousnesspredicted: no prediction
NO DIRECTIONAL PREDICTION. Whether shame is sharper before strangers or before people whose judgement you cannot escape is genuinely open, and either answer is informative.
Extraversion
- Friendlinesspredicted: drops sharply
Warmth toward a target is almost definitionally target-dependent.
- Gregariousnesspredicted: tapers
A crowd of strangers and a room of friends are plainly different experiences, and standard items rarely disambiguate.
Agreeableness
- Trustpredicted: drops sharply
An affective stance rather than a rule. Expected to gate sharply on closeness.
- Moralitypredicted: no changecontrol
A conduct rule, applied uniformly. If this gates as much as Altruism, H1 is false. A control exists to be wrong. If this one gates as steeply as the facets above it, the hypothesis is false and the site will say so.
- Altruismpredicted: drops sharply
The facet that motivated the whole instrument: IPIP items ask about people in general.
- Cooperationpredicted: no changecontrol
Conflict conduct rather than affect. A control exists to be wrong. If this one gates as steeply as the facets above it, the hypothesis is false and the site will say so.
- Sympathypredicted: drops sharply
Vicarious feeling, which plausibly requires a relationship to activate at all.
Conscientiousness
- Dutifulnesspredicted: no changecontrol
Obligation as a rule about oneself rather than a response to a person. A control exists to be wrong. If this one gates as steeply as the facets above it, the hypothesis is false and the site will say so.
The other 20: measured only as a single number
No closeness items were written for these, so this instrument makes no claim about them either way. That is a limit of what was built, not a finding that they hold steady — Orderliness may well look different at home than at work, and nothing here would show it.
Still gathering data
Everything above describes a population. This describes the evidence — how wide the error bars are, how fast they narrow, and which questions cannot be answered yet at all. Most of it is drawable with nobody in the sample, because the width of a confidence interval at a given sample size is a fact about arithmetic rather than about us.
Everything below this line is provisional. These panels describe how much the instrument can currently support, which right now is not very much: no real respondents have completed it.
They are published early on purpose. A figure that shows an empty interval today and a narrow one in a year is a far more honest account of a young instrument than a page that appears finished and turns out to be describing eleven people. Each panel says which of three things it is: real now, waiting on respondents, or waiting on something that has to be built first.
What the next respondent buys
Gathering dataSix thresholds, each quoted from a rule the site already enforces rather than chosen because it looked like a milestone. Crossing one changes what the site is allowed to say.
25 more people until "a group can be described at all".
- 25A group can be described at all — The first demographic cell clears the minimum publication size, so one slice of the population explorer can show our own respondents rather than the reference sample.The k-anonymity floor is 25. Below it nothing is published, ever.
- 100A facet mean stops being a guess — A ring mean on the 1-7 closeness scale is known to about ±0.2 scale points at 95%, which is narrower than the differences between rings the construct predicts.z·SD/√n with SD ≈ 1.1, the working ring composite SD in the scoring engine.
- 175The first hypothesis becomes testable — Enough power to detect a facet-level gradient of d = 0.3 at 80%, which is the smallest slope that would matter. Below this a null result means nothing.One-sample normal approximation at α = .05, power = .80, d = 0.30.
- 300The closeness scales get percentiles — A respondent can be told where they sit relative to other people, not just what their raw profile looks like. Below this the site shows shape and description only.Norm sampling error stops dominating measurement error around here — see the derivation in norms/ci.ts. It is a hard gate in the engine, not a UI convention.
- 1,000Slices get cut — Sex × age cells begin clearing the publication floor across the board, so group differences in the closeness measure can be published rather than only the pooled result.Twelve sex × age cells over a realistic recruitment skew, each needing 25.
- 2,000The construct survives or it does not — Enough to test incremental validity against self-concept differentiation and the moral-circle rivals — the pre-registered falsification conditions.Detecting an incremental R² of about .01 over four covariates at 80% power, which is the smallest increment worth claiming.
The maths
The floor is the k-anonymity minimum cell size, 25: below it nothing is published, ever. The percentile gate at 300 is where sampling error in the reference distribution stops dominating measurement error in the score — a hard gate in the scoring engine, not a display convention.
The testability rung is a one-sample power calculation: n = ((z1−α/2 + zpower) / d)² = ((1.96 + 0.84) / 0.30)² = 88, rounded up to leave room for incomplete closeness blocks. It is the normal approximation, which runs about two people light of the exact t-based figure.
How fast the error bars narrow
LiveTrue today, with nobody in the sample: the width of a confidence interval at a given sample size is a fact about arithmetic. Only the marker moves.
| Known to within | r = 0 | r = .30 | r = .60 |
|---|---|---|---|
| ±0.30 | 44 | 37 | 21 |
| ±0.20 | 97 | 81 | 43 |
| ±0.10 | 385 | 320 | 161 |
| ±0.05 | 1,538 | 1,274 | 633 |
The maths
Intervals on a correlation are built on Fisher’s z = atanh(r), where the sampling distribution is near-normal with SE = 1/√(n−3), then transformed back with tanh. The naïve r ± z·SE produces intervals that run past ±1, which is how you can tell nobody looked at the output.
The shape is 1/√n, and that is the whole story of small samples: 25 → 100 halves the interval, and so does 1,000 → 4,000. Precision is cheap at the start and ruinously expensive later.
What four items actually buys you
LiveReal numbers from 512,295 respondents aged 18+ from johnson's ipip-neo-120 dataset (osf.io/wxvth), scored with this engine, scored by this engine. This panel never needed our own respondents, and it is the one most other tests in this market decline to publish.
Describes: 512,295 respondents aged 18+ from Johnson's IPIP-NEO-120 dataset (osf.io/wxvth), scored with this engine. Median facet band is ±19 percentile points.
| Facet | α | SD | SEM | 80% band | in percentile points |
|---|---|---|---|---|---|
| Self-Efficacy | 0.78 | 2.47 | 1.17 | ±25 | |
| Cooperation | 0.69 | 3.51 | 1.96 | ±25 | |
| Self-Consciousness | 0.71 | 3.60 | 1.94 | ±24 | |
| Altruism | 0.73 | 2.60 | 1.35 | ±22 | |
| Adventurousness | 0.72 | 3.27 | 1.74 | ±22 | |
| Dutifulness | 0.67 | 2.63 | 1.50 | ±22 | |
| Liberalism | 0.66 | 3.53 | 2.07 | ±21 | |
| Modesty | 0.73 | 3.42 | 1.78 | ±21 | |
| Intellect | 0.74 | 3.54 | 1.80 | ±20 | |
| Artistic Interests | 0.74 | 3.59 | 1.82 | ±19 | |
| Sympathy | 0.72 | 3.09 | 1.63 | ±19 | |
| Cheerfulness | 0.80 | 3.23 | 1.43 | ±19 | |
| Vulnerability | 0.77 | 3.63 | 1.73 | ±19 | |
| Assertiveness | 0.85 | 3.45 | 1.33 | ±19 | |
| Emotionality | 0.65 | 3.00 | 1.76 | ±19 | |
| Morality | 0.74 | 2.98 | 1.51 | ±19 | |
| Anxiety | 0.79 | 3.78 | 1.75 | ±18 | |
| Activity Level | 0.70 | 3.15 | 1.73 | ±18 | |
| Orderliness | 0.84 | 4.32 | 1.73 | ±18 | |
| Gregariousness | 0.79 | 4.00 | 1.83 | ±17 | |
| Self-Discipline | 0.72 | 3.21 | 1.70 | ±17 | |
| Achievement-Striving | 0.79 | 3.23 | 1.48 | ±17 | |
| Friendliness | 0.81 | 3.59 | 1.56 | ±16 | |
| Immoderation | 0.71 | 3.42 | 1.84 | ±16 | |
| Imagination | 0.75 | 3.40 | 1.71 | ±16 | |
| Excitement-Seeking | 0.73 | 3.34 | 1.73 | ±16 | |
| Depression | 0.85 | 3.87 | 1.51 | ±14 | |
| Trust | 0.86 | 3.52 | 1.34 | ±14 | |
| Anger | 0.87 | 4.13 | 1.52 | ±12 | |
| Cautiousness | 0.88 | 4.11 | 1.41 | ±12 |
The maths
SEM = SD·√(1−α) is how much a score would move on a retest. The 80% band is ±1.2816 × SE, and the band in percentile points comes from reading those raw limits off the exact discrete CDF rather than assuming a normal shape the 17-point support does not have.
Evaluated at each facet’s median, where percentile bands are widest, because that is the typical reader’s case. Computed by the same function the results page calls — two implementations of a confidence interval eventually disagree, and the one that disagrees on a public methods page is the expensive one.
Three more panels — the four pre-registered hypotheses with what would falsify each, the single unmeasured number the person-level index rests on, and the eight instruments that are not in the test — are on the analytics page.
What these numbers are not
- Not a probability sample. Everyone here chose to take a personality test on the internet. That is a real population and a useful one, and it is not the general public. Treat differences between groups as descriptions of who took this, not estimates of who exists.
- Not stable at small n. Groups below 25 people are never published, and groups just above it move a lot as new people arrive. The interval on each average is the honest width of what is known.
- Not individually identifying, by construction. Counts are rounded to the nearest 5, at most 3 filters can be combined, a cell is hidden when publishing it would let a neighbouring cell be worked out by subtraction, and a cell is not refreshed until enough new people have arrived that the refresh cannot reveal who they were.
- Not the gradient module. Everything on this page is the validated Big Five core. The closeness measure has no reference sample yet, so it has no population view — it appears here when there is a population to describe.
Methods and formulas: /methods. Raw data, codebook and citation: /data.