The evidence
How much do we know yet
Every other page here shows what the instrument measured. This one shows how much of it can be believed: the width of each error bar, the rate at which more people narrow it, the thresholds at which the site becomes allowed to say new things, and the four pre-registered hypotheses — three of which no number of respondents can currently test.
It is published now, while the answer is mostly “not yet”, because a young instrument that only shows finished-looking pages is asking to be trusted on the strength of its typography. Most of what is here does not need a single respondent to be true. The width of a confidence interval at four hundred people is a fact about arithmetic; only the marker saying where we are moves.
Everything below this line is provisional. These panels describe how much the instrument can currently support, which right now is not very much: no real respondents have completed it.
They are published early on purpose. A figure that shows an empty interval today and a narrow one in a year is a far more honest account of a young instrument than a page that appears finished and turns out to be describing eleven people. Each panel says which of three things it is: real now, waiting on respondents, or waiting on something that has to be built first.
What the next respondent buys
Gathering dataSix thresholds, each quoted from a rule the site already enforces rather than chosen because it looked like a milestone. Crossing one changes what the site is allowed to say.
25 more people until "a group can be described at all".
- 25A group can be described at all — The first demographic cell clears the minimum publication size, so one slice of the population explorer can show our own respondents rather than the reference sample.The k-anonymity floor is 25. Below it nothing is published, ever.
- 100A facet mean stops being a guess — A ring mean on the 1-7 closeness scale is known to about ±0.2 scale points at 95%, which is narrower than the differences between rings the construct predicts.z·SD/√n with SD ≈ 1.1, the working ring composite SD in the scoring engine.
- 175The first hypothesis becomes testable — Enough power to detect a facet-level gradient of d = 0.3 at 80%, which is the smallest slope that would matter. Below this a null result means nothing.One-sample normal approximation at α = .05, power = .80, d = 0.30.
- 300The closeness scales get percentiles — A respondent can be told where they sit relative to other people, not just what their raw profile looks like. Below this the site shows shape and description only.Norm sampling error stops dominating measurement error around here — see the derivation in norms/ci.ts. It is a hard gate in the engine, not a UI convention.
- 1,000Slices get cut — Sex × age cells begin clearing the publication floor across the board, so group differences in the closeness measure can be published rather than only the pooled result.Twelve sex × age cells over a realistic recruitment skew, each needing 25.
- 2,000The construct survives or it does not — Enough to test incremental validity against self-concept differentiation and the moral-circle rivals — the pre-registered falsification conditions.Detecting an incremental R² of about .01 over four covariates at 80% power, which is the smallest increment worth claiming.
The maths
The floor is the k-anonymity minimum cell size, 25: below it nothing is published, ever. The percentile gate at 300 is where sampling error in the reference distribution stops dominating measurement error in the score — a hard gate in the scoring engine, not a display convention.
The testability rung is a one-sample power calculation: n = ((z1−α/2 + zpower) / d)² = ((1.96 + 0.84) / 0.30)² = 88, rounded up to leave room for incomplete closeness blocks. It is the normal approximation, which runs about two people light of the exact t-based figure.
How fast the error bars narrow
LiveTrue today, with nobody in the sample: the width of a confidence interval at a given sample size is a fact about arithmetic. Only the marker moves.
| Known to within | r = 0 | r = .30 | r = .60 |
|---|---|---|---|
| ±0.30 | 44 | 37 | 21 |
| ±0.20 | 97 | 81 | 43 |
| ±0.10 | 385 | 320 | 161 |
| ±0.05 | 1,538 | 1,274 | 633 |
The maths
Intervals on a correlation are built on Fisher’s z = atanh(r), where the sampling distribution is near-normal with SE = 1/√(n−3), then transformed back with tanh. The naïve r ± z·SE produces intervals that run past ±1, which is how you can tell nobody looked at the output.
The shape is 1/√n, and that is the whole story of small samples: 25 → 100 halves the interval, and so does 1,000 → 4,000. Precision is cheap at the start and ruinously expensive later.
What four items actually buys you
LiveReal numbers from 512,295 respondents aged 18+ from johnson's ipip-neo-120 dataset (osf.io/wxvth), scored with this engine, scored by this engine. This panel never needed our own respondents, and it is the one most other tests in this market decline to publish.
Describes: 512,295 respondents aged 18+ from Johnson's IPIP-NEO-120 dataset (osf.io/wxvth), scored with this engine. Median facet band is ±19 percentile points.
| Facet | α | SD | SEM | 80% band | in percentile points |
|---|---|---|---|---|---|
| Self-Efficacy | 0.78 | 2.47 | 1.17 | ±25 | |
| Cooperation | 0.69 | 3.51 | 1.96 | ±25 | |
| Self-Consciousness | 0.71 | 3.60 | 1.94 | ±24 | |
| Altruism | 0.73 | 2.60 | 1.35 | ±22 | |
| Adventurousness | 0.72 | 3.27 | 1.74 | ±22 | |
| Dutifulness | 0.67 | 2.63 | 1.50 | ±22 | |
| Liberalism | 0.66 | 3.53 | 2.07 | ±21 | |
| Modesty | 0.73 | 3.42 | 1.78 | ±21 | |
| Intellect | 0.74 | 3.54 | 1.80 | ±20 | |
| Artistic Interests | 0.74 | 3.59 | 1.82 | ±19 | |
| Sympathy | 0.72 | 3.09 | 1.63 | ±19 | |
| Cheerfulness | 0.80 | 3.23 | 1.43 | ±19 | |
| Vulnerability | 0.77 | 3.63 | 1.73 | ±19 | |
| Assertiveness | 0.85 | 3.45 | 1.33 | ±19 | |
| Emotionality | 0.65 | 3.00 | 1.76 | ±19 | |
| Morality | 0.74 | 2.98 | 1.51 | ±19 | |
| Anxiety | 0.79 | 3.78 | 1.75 | ±18 | |
| Activity Level | 0.70 | 3.15 | 1.73 | ±18 | |
| Orderliness | 0.84 | 4.32 | 1.73 | ±18 | |
| Gregariousness | 0.79 | 4.00 | 1.83 | ±17 | |
| Self-Discipline | 0.72 | 3.21 | 1.70 | ±17 | |
| Achievement-Striving | 0.79 | 3.23 | 1.48 | ±17 | |
| Friendliness | 0.81 | 3.59 | 1.56 | ±16 | |
| Immoderation | 0.71 | 3.42 | 1.84 | ±16 | |
| Imagination | 0.75 | 3.40 | 1.71 | ±16 | |
| Excitement-Seeking | 0.73 | 3.34 | 1.73 | ±16 | |
| Depression | 0.85 | 3.87 | 1.51 | ±14 | |
| Trust | 0.86 | 3.52 | 1.34 | ±14 | |
| Anger | 0.87 | 4.13 | 1.52 | ±12 | |
| Cautiousness | 0.88 | 4.11 | 1.41 | ±12 |
The maths
SEM = SD·√(1−α) is how much a score would move on a retest. The 80% band is ±1.2816 × SE, and the band in percentile points comes from reading those raw limits off the exact discrete CDF rather than assuming a normal shape the 17-point support does not have.
Evaluated at each facet’s median, where percentile bands are widest, because that is the typical reader’s case. Computed by the same function the results page calls — two implementations of a confidence interval eventually disagree, and the one that disagrees on a public methods page is the expensive one.
The four hypotheses, and what would kill each one
Gathering dataPre-registered before any data existed. Three of the four cannot be tested by any number of respondents, because the instruments they need are not in the item bank — and saying so is the difference between a pre-registration and a decoration.
- H1Facets differ in how much they gate: conduct facets stay flat across closeness, affective facets slope.
Falsified by: The three control facets — Morality, Cooperation, Modesty — sloping as steeply as Altruism, Sympathy and Trust. That is a single contrast within one domain, so it cannot be rescued by an appeal to some other trait behaving differently.
Gathering dataCompleted gradient blocks from enough respondents to detect d = 0.3.0 of 1750% - H2Gradient magnitude is a stable, trait-like individual difference.
Falsified by: A test-retest correlation on the Social Contraction Index that is not distinguishable from the retest of a noise variable. If it is not stable it is not a trait, and individual-level feedback should not ship at all.
BlockedA retest wave: the same people, several weeks apart.More respondents will not help. No retest mechanism exists yet. A retrieval code can bring somebody back, but nothing invites them to, and nothing links the two sessions.
- H3Large gradients on affective facets have developmental correlates — attachment history, early social environment, learned protective gating.
Falsified by: No association between gradient magnitude and the early-environment block.
BlockedThe separately-consented sensitive block.More respondents will not help. The early-environment block is specified but not written into the item bank.
- H4Gradient shape predicts things the mean score does not.
Falsified by: No incremental validity over global trait level, self-concept differentiation, neuroticism, and the MESx + IOS moral-circle measures. All four rivals are administered specifically so this can be tested rather than asserted.
BlockedThe rival measures, plus outcomes to predict.More respondents will not help. None of the four rival measures is in the instrument yet.
The maths
H1’s threshold is the power calculation above at d = 0.30. H4’s is the sample needed to detect an incremental R² ≈ .01 over four covariates at 80% power, which is the smallest increment worth claiming for a new construct.
H2 has no sample size at all, and that is not an omission. Stability is a test–retest question: the same people, twice, weeks apart. No amount of new respondents substitutes for one returning one.
The person-level index rests on one unmeasured number
Not collected yetThe Social Contraction Index averages standardised per-facet slopes. Whether that average is a usable measurement depends entirely on how much the slopes agree with each other — a quantity nobody has ever estimated, in anyone.
The maths
A composite of k parts correlating r̄ on average has reliability k·r̄ / (1 + (k−1)·r̄) — Spearman-Brown, with the mean inter-part correlation standing in for single-part reliability.
At k = 9 and the assumed r̄ = 0.25 that gives α ≈ 0.75, which is usable. At r̄ = .10 it gives α ≈ 0.50, at which point individual-level closeness feedback should not ship at all.
Both numbers are computed from the same formula the scoring engine uses, and r̄ = 0.25 is an assumption written into the code with a comment saying it is untested. This panel exists so that assumption is visible from outside the repository.
Eight instruments that are not in the test
Not collected yetIncluding all three rival measures the falsification conditions name, and the ring-anchor calibration that would tell a flat gradient apart from a compressed reading of the rings. Adding every one would make the instrument 66 items longer.
None of these is in the instrument. Each is written out because the reason it is missing matters: three of them are the rival explanations this project pre-registered as the conditions under which its own construct fails, and one of them — ring anchor calibration — is the reason a flat gradient is currently ambiguous.
Adopting all eight would add 66 items to a 240-item instrument — about a quarter longer, with the dropout that implies. Statistical power bought with fewer respondents is not obviously a good trade, which is why this is a list rather than a plan.
Self-concept differentiation
+20 itemsHow much your self-description changes between social roles.
What it would settle: The single most dangerous rival explanation. If the closeness gradient is just SCD wearing different clothes, the construct is a reframing and we publish that.
How it would be scored
Donahue et al. (1993): rate the same trait adjectives in several roles, then take 1 − (variance of the first principal component ÷ total variance) across the role × trait matrix. High SCD means the roles disagree.
Moral Expansiveness Scale (short form)
+15 itemsHow far out your circle of moral concern reaches.
What it would settle: Whether the gradient is a behavioural fact or a restatement of moral-circle breadth.
How it would be scored
Crimston et al. (2018): entities are placed in one of four concentric circles, scored 3/2/1/0 and summed. A direct competitor to the gradient — it is the same concentric picture drawn for moral standing rather than for behaviour.
Inclusion of Other in the Self
+4 itemsFelt closeness, as overlapping circles, one item per ring.
What it would settle: THE RING ANCHOR PROBLEM. "Acquaintances" is a word, and one person's acquaintance is another's friend. Without a calibration item, a shallow gradient and a compressed interpretation of the rings are indistinguishable.
How it would be scored
Aron, Aron & Smollan (1992): a single pictorial 1–7 item per target. Four of them calibrate what each ring actually means to this respondent.
Expression cost
+12 itemsWhat it costs you to express a trait toward each ring, as distinct from whether you do.
What it would settle: Whether a flat profile means "I am the same with everyone" or "I work hard to be".
How it would be scored
Parallel 1–7 stems per ring — effortful/automatic — scored as a second surface over the same facets. The contrast of interest is cost against expression: high expression at high cost is a different person from high expression at no cost.
Circle architecture
+4 itemsHow many people are actually in each ring.
What it would settle: The single biggest limitation of the current design. Our rings are ORDINAL, which is why the engine uses orthogonal polynomial contrasts and refuses to fit a discounting k. With counts, a real discount function becomes estimable.
How it would be scored
Free counts per ring, log-transformed. Tests the Dunbar-layer prediction that ring sizes scale by roughly a factor of three, and lets the gradient be re-expressed against a ratio-scaled x-axis instead of four ordinal categories.
Switch or dial
+3 itemsWhether the change is continuous or has a boundary.
What it would settle: Whether "cliff" is a real shape class or an artefact of four measurement points.
How it would be scored
A forced choice plus a boundary-placement item, checked against the cubic contrast the engine already computes as a noise probe. Two independent routes to the same answer is what makes either believable.
Motive attribution
+4 itemsWhy you think you withhold — protection, capacity, indifference, or reciprocity.
What it would settle: The developmental hypothesis. Protective gating and depleted capacity produce the same curve and mean entirely different things.
How it would be scored
Four 1–7 stems, one per motive, per respondent rather than per facet.
Attention and consistency checks
+4 itemsWhether the answers were given by somebody reading the questions.
What it would settle: Whether any of the above can be believed. A falling median quality score is the earliest signal of coordinated data poisoning — which is worse than an outage, because an outage ends and a corrupted dataset does not.
How it would be scored
Instructed-response items, a long-string run detector, and the within-stem agreement the gradient scorer already computes. Combined into a single quality score per session.
The maths
Each entry states its own scoring rule — the SCD variance ratio, the MESx 3/2/1/0 circle sum, the IOS single pictorial item per ring — in enough detail to be checked against the source paper before anybody writes an item.
The one with the largest methodological consequence is circle architecture. Our four rings are ordinal categories, which is why the engine decomposes them with orthogonal polynomial contrasts and refuses to fit a hyperbolic discount rate: k shifts by orders of magnitude depending on an arbitrary ring-to-N coding. With real ring counts, the x-axis becomes ratio-scaled and a genuine discount function becomes estimable.
What this page is not
- Not site traffic. Nothing here counts visitors, sessions, referrers or anything else about who reads the site. It counts completed assessments and what they support statistically. Web analytics on this site are cookieless and aggregate, and never appear here.
- Not a results page. No figure here describes any individual, and none can be filtered toward one. The same k-anonymity rules that govern the population explorer govern this.
- Not a promise. A threshold says what becomes statistically possible at a sample size, not that the sample will be reached or that the hypothesis will survive it. The falsification conditions are published in advance precisely so they cannot be quietly moved later.
The population itself: /explore. Formulas and scoring: /methods. The construct and its hypotheses: /preprint.