Concentric

The population

Explore the data

Everything an individual result compares you against, shown directly. Same items, same scoring engine, read across everyone rather than one person. Each figure carries the number of people behind it and the date it describes, because a statistic without either is a decoration.

No aggregate release has been published yet. Every distribution, small multiple and age trajectory below is drawn from the reference sample and is real. The three sections that describe our own respondents — group comparison, the correlation matrix, and the closeness coverage counts — appear once enough people have taken the assessment for a cell to clear the minimum publication size.

Showing everyone. You can combine up to 3 filters — beyond that, a group can be small enough to identify someone even when the count looks large.

How Neuroticism is distributed

41220lowerhigher
mean 11.18SD 2.64α 0.90n 512,285

The steps are real. A facet score is the sum of four whole-numbered answers, so only certain values are possible and a smooth curve would be inventing shape the data does not have. Height is the share of people at each score, not a count, so a filtered group stays comparable with the whole population behind it.

Neuroticism across groups

Averages with their intervals. Sex and age differences in personality are real and small — usually a fraction of the spread within any one group — so the intervals are the point, not the dots.

InfinityNaN-Infinity

Each bar is a 95% interval around that group’s average. Where two intervals overlap, the difference between those groups is not established by this data — however far apart the dots look. The axis is zoomed to the range the averages occupy, which is a small slice of the full 420 scale.

All thirty narrow traits

Every facet on one screen, drawn on a shared scale so the shapes can be compared directly rather than one at a time from memory. Select any of them to read what it measures.

Neuroticism across the age bands

The whole distribution for each age group, not just its average. Two groups can share a mean and be shaped completely differently — one clustered tightly, one split in two — and a table of averages hides exactly that. Every curve is drawn on the same axis and the same vertical scale, so both position and spread can be read across rows.

share of each group · peak 2.7%
18-19
20-21
22-25
26-30
31-39
40+
41220
Distribution of Neuroticism by age band. Every curve shares one horizontal axis and one vertical scale, so both position and spread are comparable between rows. Describes: 512,295 respondents aged 18+ from Johnson's IPIP-NEO-120 dataset (osf.io/wxvth), scored with this engine
Show the numbers
age bandnmeanSDα
18-19107,07211.492.510.89
20-21100,87611.172.560.90
22-2599,22811.272.660.90
26-3068,08811.272.730.91
31-3968,45111.122.740.91
40+68,57010.552.660.90

The five domains by age

Where each domain sits in each age band, with the interval around every average. The intervals are the point: when they overlap, the difference between two bands is not something this sample can distinguish, however suggestive the line between them looks.

score range 4–20 · 95% interval
16.613.19.6
18-1920-2122-2526-3031-3940+
NeuroticismExtraversionOpennessAgreeablenessConscientiousness
Group averages by age band, with 95% intervals. These are different people, not the same people followed over time — a difference between bands may reflect generational differences as much as change with age. Describes: 512,295 respondents aged 18+ from Johnson's IPIP-NEO-120 dataset (osf.io/wxvth), scored with this engine
Show the numbers
Trait18-1920-2122-2526-3031-3940+
Neuroticism11.49 ±0.0211.17 ±0.0211.27 ±0.0211.27 ±0.0211.12 ±0.0210.55 ±0.02
Extraversion13.87 ±0.0113.93 ±0.0113.70 ±0.0113.50 ±0.0213.39 ±0.0213.27 ±0.02
Openness13.67 ±0.0113.63 ±0.0114.00 ±0.0114.11 ±0.0213.96 ±0.0213.91 ±0.02
Agreeableness14.61 ±0.0114.67 ±0.0114.65 ±0.0114.76 ±0.0214.97 ±0.0115.51 ±0.01
Conscientiousness14.08 ±0.0114.56 ±0.0114.58 ±0.0114.82 ±0.0215.10 ±0.0215.62 ±0.02

These are different people measured once, not the same people followed over time. A difference between two bands may be change that comes with age, or it may be that people born thirty years apart answer differently — a single sweep cannot tell those apart, and nothing on this page should be read as a claim that anyone changes.

Where women and men differ

Every facet, sorted by the size of the difference rather than by whether one exists. Differences are shown in standard deviations, because a gap of half a point means something different on a narrow trait than on a wide one and these thirty scales have to be comparable to each other.

Cohen's d · positive favours women
Emotionality
Sympathy
Anxiety
Vulnerability
Altruism
Modesty
Morality
Artistic Interests
Cooperation
Achievement-Striving
Activity Level
Dutifulness
Anger
Depression
Immoderation
Cheerfulness
Liberalism
Orderliness
Self-Discipline
Self-Consciousness
Friendliness
Gregariousness
Trust
Self-Efficacy
Imagination
Cautiousness
Adventurousness
Assertiveness
Excitement-Seeking
Intellect
0.65+0.65
Standardised difference between men and women on each scale, sorted by size, with 95% intervals. Bars to the right mean women score higher. The shaded strip is the range where a difference is too small to be worth interpreting. Describes: 512,295 respondents aged 18+ from Johnson's IPIP-NEO-120 dataset (osf.io/wxvth), scored with this engine
Show the numbers
Scaled95% low95% highn
Emotionality0.6140.6080.620512,080
Sympathy0.5410.5350.547512,133
Anxiety0.5350.5300.541512,223
Vulnerability0.5210.5150.527512,097
Altruism0.5030.4980.509512,141
Modesty0.4500.4440.456512,113
Morality0.3850.3800.391512,075
Artistic Interests0.3130.3070.318512,113
Cooperation0.2900.2840.295512,074
Achievement-Striving0.2380.2330.244512,054
Activity Level0.2370.2310.242512,075
Dutifulness0.1940.1890.200512,051
Anger0.1890.1830.194512,162
Depression0.1750.1690.181512,101
Immoderation0.1600.1550.166512,044
Cheerfulness0.1310.1250.136512,012
Liberalism0.1200.1150.126511,998
Orderliness0.0980.0930.104512,124
Self-Discipline0.0970.0920.103512,080
Self-Consciousness0.0920.0870.098512,083
Friendliness0.0890.0830.094512,193
Gregariousness0.0740.0690.080512,124
Trust0.0200.0140.026512,185
Self-Efficacy-0.026-0.031-0.020512,137
Imagination-0.061-0.067-0.055512,178
Cautiousness-0.064-0.069-0.058512,155
Adventurousness-0.113-0.119-0.107512,103
Assertiveness-0.142-0.147-0.136512,080
Excitement-Seeking-0.205-0.211-0.200512,032
Intellect-0.230-0.236-0.225512,046

With a large enough sample almost every difference becomes detectable, which is why the shaded strip matters more than the bars: inside it, a difference is real in the sense that it is not zero, and too small to tell you anything about a person in front of you. Group averages describe groups. The overlap between two of these distributions is always far larger than the gap between their centres.

How the traits relate

Every facet against every other facet. This is the structure the five broad domains are a summary of — and the places where the summary fails are visible here and nowhere else on the site.

The correlation matrix is computed over our own respondents, and no aggregate release has been published yet.

Nothing is hidden to make the page look tidier. Try removing a filter.

Where a single number is doing the most work

Everything above this point is the conventional Big Five: one number per trait, with an interval around it. That interval covers measurement error — how much the number would move if you answered the same statements again. It does not cover the other thing that moves it, which is who the statements are about.

The public-domain items behind several of these facets ask about people in general. “I go out of my way to help others.” A person whose answer is alwaysfor the four people closest to them and rarely for anyone else has no honest way to answer, and the instrument records the average of two behaviours rather than either one. The number is a reading taken at one unstated social distance and reported as though distance were irrelevant — so its interval is too narrow, and not by a rounding amount.

That error also does not wash out when six facets are averaged into a domain. Measurement error is roughly independent between facets, so it shrinks as they combine. Withholding is not independent: the whole premise is that it is a property of the person, so someone who reserves Altruism for their inner circle plausibly reserves Sympathy too. Correlated components survive averaging, which is how a facet-level problem becomes a domain-level one.

How much of each domain carries a closeness measurement

Coverage, not magnitude. A filled cell means closeness statements were written for that facet, so the question can be asked of it. Nobody has answered them yet, so a domain can be five-sixths covered and still turn out perfectly steady — which is what the controls below are for.

Agreeableness
5 of 6
Trust
Morality
Altruism
Cooperation
Modesty
Sympathy
Neuroticism
2 of 6
Anxiety
Anger
Depression
Self-Consciousness
Immoderation
Vulnerability
Extraversion
2 of 6
Friendliness
Gregariousness
Assertiveness
Activity Level
Excitement-Seeking
Cheerfulness
Conscientiousness
1 of 6
Self-Efficacy
Orderliness
Dutifulness
Achievement-Striving
Self-Discipline
Cautiousness
Openness
0 of 6
Imagination
Artistic Interests
Emotionality
Adventurousness
Intellect
Liberalism

Agreeableness is the domain most exposed, at five facets of six — which is unsurprising, because it is the domain that motivated building any of this. Openness is not covered at all. And the interesting design detail is that Agreeableness is exposed in both directions: three of its five are predicted to gate steeply, and two are controls predicted not to move. If the controls gate as hard as the rest, the hypothesis is wrong, and it can be tested inside a single domain without comparing across domains at all.

The ten facets, and what was predicted for each

Nothing below has been measured yet. These are the predictions the instrument was built to test, registered before any responses were collected. They are here so that the record of what we expected cannot be quietly revised once the data arrives.

Neuroticism

  • Angerpredicted: risesreversal control

    REVERSAL CONTROL, and the strongest available defence against response-style artifacts. Predicted to be expressed MORE toward close others -- you are short with family and civil with strangers -- consistent with Clifton (2014). A uniform response bias cannot produce opposite-signed slopes. A control exists to be wrong. If this one falls with distance like the rest, the pattern is a response habit rather than a finding.

  • Self-Consciousnesspredicted: no prediction

    NO DIRECTIONAL PREDICTION. Whether shame is sharper before strangers or before people whose judgement you cannot escape is genuinely open, and either answer is informative.

Extraversion

  • Friendlinesspredicted: drops sharply

    Warmth toward a target is almost definitionally target-dependent.

  • Gregariousnesspredicted: tapers

    A crowd of strangers and a room of friends are plainly different experiences, and standard items rarely disambiguate.

Agreeableness

  • Trustpredicted: drops sharply

    An affective stance rather than a rule. Expected to gate sharply on closeness.

  • Moralitypredicted: no changecontrol

    A conduct rule, applied uniformly. If this gates as much as Altruism, H1 is false. A control exists to be wrong. If this one gates as steeply as the facets above it, the hypothesis is false and the site will say so.

  • Altruismpredicted: drops sharply

    The facet that motivated the whole instrument: IPIP items ask about people in general.

  • Cooperationpredicted: no changecontrol

    Conflict conduct rather than affect. A control exists to be wrong. If this one gates as steeply as the facets above it, the hypothesis is false and the site will say so.

  • Sympathypredicted: drops sharply

    Vicarious feeling, which plausibly requires a relationship to activate at all.

Conscientiousness

  • Dutifulnesspredicted: no changecontrol

    Obligation as a rule about oneself rather than a response to a person. A control exists to be wrong. If this one gates as steeply as the facets above it, the hypothesis is false and the site will say so.

The other 20: measured only as a single number

No closeness items were written for these, so this instrument makes no claim about them either way. That is a limit of what was built, not a finding that they hold steady — Orderliness may well look different at home than at work, and nothing here would show it.

AnxietyDepressionImmoderationVulnerabilityAssertivenessActivity LevelExcitement-SeekingCheerfulnessImaginationArtistic InterestsEmotionalityAdventurousnessIntellectLiberalismModestySelf-EfficacyOrderlinessAchievement-StrivingSelf-DisciplineCautiousness

Still gathering data

Everything above describes a population. This describes the evidence — how wide the error bars are, how fast they narrow, and which questions cannot be answered yet at all. Most of it is drawable with nobody in the sample, because the width of a confidence interval at a given sample size is a fact about arithmetic rather than about us.

Everything below this line is provisional. These panels describe how much the instrument can currently support, which right now is not very much: no real respondents have completed it.

They are published early on purpose. A figure that shows an empty interval today and a narrow one in a year is a far more honest account of a young instrument than a page that appears finished and turns out to be describing eleven people. Each panel says which of three things it is: real now, waiting on respondents, or waiting on something that has to be built first.

What the next respondent buys

Gathering data

Six thresholds, each quoted from a rule the site already enforces rather than chosen because it looked like a milestone. Crossing one changes what the site is allowed to say.

251001753001,0002,000

25 more people until "a group can be described at all".

  1. 25A group can be described at allThe first demographic cell clears the minimum publication size, so one slice of the population explorer can show our own respondents rather than the reference sample.The k-anonymity floor is 25. Below it nothing is published, ever.
  2. 100A facet mean stops being a guessA ring mean on the 1-7 closeness scale is known to about ±0.2 scale points at 95%, which is narrower than the differences between rings the construct predicts.z·SD/√n with SD ≈ 1.1, the working ring composite SD in the scoring engine.
  3. 175The first hypothesis becomes testableEnough power to detect a facet-level gradient of d = 0.3 at 80%, which is the smallest slope that would matter. Below this a null result means nothing.One-sample normal approximation at α = .05, power = .80, d = 0.30.
  4. 300The closeness scales get percentilesA respondent can be told where they sit relative to other people, not just what their raw profile looks like. Below this the site shows shape and description only.Norm sampling error stops dominating measurement error around here — see the derivation in norms/ci.ts. It is a hard gate in the engine, not a UI convention.
  5. 1,000Slices get cutSex × age cells begin clearing the publication floor across the board, so group differences in the closeness measure can be published rather than only the pooled result.Twelve sex × age cells over a realistic recruitment skew, each needing 25.
  6. 2,000The construct survives or it does notEnough to test incremental validity against self-concept differentiation and the moral-circle rivals — the pre-registered falsification conditions.Detecting an incremental R² of about .01 over four covariates at 80% power, which is the smallest increment worth claiming.
The maths

The floor is the k-anonymity minimum cell size, 25: below it nothing is published, ever. The percentile gate at 300 is where sampling error in the reference distribution stops dominating measurement error in the score — a hard gate in the scoring engine, not a display convention.

The testability rung is a one-sample power calculation: n = ((z1−α/2 + zpower) / d)² = ((1.96 + 0.84) / 0.30)² = 88, rounded up to leave room for incomplete closeness blocks. It is the normal approximation, which runs about two people light of the exact t-based figure.

How fast the error bars narrow

Live

True today, with nobody in the sample: the width of a confidence interval at a given sample size is a fact about arithmetic. Only the marker moves.

±0.1±0.2±0.3±0.4r = 0r = .30r = .6020501003001,0003,00010,000respondents (logarithmic)
Sample size needed for a correlation to be known to within a given half-width.
Known to withinr = 0r = .30r = .60
±0.30443721
±0.20978143
±0.10385320161
±0.051,5381,274633
The maths

Intervals on a correlation are built on Fisher’s z = atanh(r), where the sampling distribution is near-normal with SE = 1/√(n−3), then transformed back with tanh. The naïve r ± z·SE produces intervals that run past ±1, which is how you can tell nobody looked at the output.

The shape is 1/√n, and that is the whole story of small samples: 25 → 100 halves the interval, and so does 1,000 → 4,000. Precision is cheap at the start and ruinously expensive later.

What four items actually buys you

Live

Real numbers from 512,295 respondents aged 18+ from johnson's ipip-neo-120 dataset (osf.io/wxvth), scored with this engine, scored by this engine. This panel never needed our own respondents, and it is the one most other tests in this market decline to publish.

Describes: 512,295 respondents aged 18+ from Johnson's IPIP-NEO-120 dataset (osf.io/wxvth), scored with this engine. Median facet band is ±19 percentile points.

FacetαSDSEM80% bandin percentile points
Self-Efficacy0.782.471.17±25
Cooperation0.693.511.96±25
Self-Consciousness0.713.601.94±24
Altruism0.732.601.35±22
Adventurousness0.723.271.74±22
Dutifulness0.672.631.50±22
Liberalism0.663.532.07±21
Modesty0.733.421.78±21
Intellect0.743.541.80±20
Artistic Interests0.743.591.82±19
Sympathy0.723.091.63±19
Cheerfulness0.803.231.43±19
Vulnerability0.773.631.73±19
Assertiveness0.853.451.33±19
Emotionality0.653.001.76±19
Morality0.742.981.51±19
Anxiety0.793.781.75±18
Activity Level0.703.151.73±18
Orderliness0.844.321.73±18
Gregariousness0.794.001.83±17
Self-Discipline0.723.211.70±17
Achievement-Striving0.793.231.48±17
Friendliness0.813.591.56±16
Immoderation0.713.421.84±16
Imagination0.753.401.71±16
Excitement-Seeking0.733.341.73±16
Depression0.853.871.51±14
Trust0.863.521.34±14
Anger0.874.131.52±12
Cautiousness0.884.111.41±12
The maths

SEM = SD·√(1−α) is how much a score would move on a retest. The 80% band is ±1.2816 × SE, and the band in percentile points comes from reading those raw limits off the exact discrete CDF rather than assuming a normal shape the 17-point support does not have.

Evaluated at each facet’s median, where percentile bands are widest, because that is the typical reader’s case. Computed by the same function the results page calls — two implementations of a confidence interval eventually disagree, and the one that disagrees on a public methods page is the expensive one.

Three more panels — the four pre-registered hypotheses with what would falsify each, the single unmeasured number the person-level index rests on, and the eight instruments that are not in the test — are on the analytics page.

What these numbers are not

  • Not a probability sample. Everyone here chose to take a personality test on the internet. That is a real population and a useful one, and it is not the general public. Treat differences between groups as descriptions of who took this, not estimates of who exists.
  • Not stable at small n. Groups below 25 people are never published, and groups just above it move a lot as new people arrive. The interval on each average is the honest width of what is known.
  • Not individually identifying, by construction. Counts are rounded to the nearest 5, at most 3 filters can be combined, a cell is hidden when publishing it would let a neighbouring cell be worked out by subtraction, and a cell is not refreshed until enough new people have arrived that the refresh cannot reveal who they were.
  • Not the gradient module. Everything on this page is the validated Big Five core. The closeness measure has no reference sample yet, so it has no population view — it appears here when there is a population to describe.

Methods and formulas: /methods. Raw data, codebook and citation: /data.