Concentric

Preprint

Social Proximity Conditionality

Measuring Big Five trait expression as a function of closeness to the target. Written before data collection, and published in that state deliberately.

Written before data collection. Every quantity in this paper is a design target or a prediction, not a result. It is published in this state so that the record of what was expected cannot be revised after the fact. Nothing here has been peer reviewed.

Social Proximity Conditionality: Measuring Big Five Trait Expression as a Function of Closeness to the Target

Robert Ticknor

Preprint — Version 1.0. Written August 2026, before data collection. Documentation licensed CC BY 4.0. Instrument items and data: CC0 1.0.


Abstract

Standard Big Five instruments return one score per facet: a dispositional point estimate that presumes trait expression is reasonably consistent across social contexts. For a substantial minority of respondents that presumption fails, and the instruments contain no mechanism to detect the failure. The case is sharpest for the Agreeableness facets Altruism and Sympathy, whose items reference people in general — the needy, strangers, humanity at large. A respondent whose prosocial response is strongly conditional on closeness to the target has no accurate way to answer: any single response either overstates the value toward distant others or understates it toward close others. The instrument returns approximately the mean of a bimodal distribution the respondent does not possess.

I propose measuring the gradient rather than the point. Selected facets are administered at four ordered levels of closeness — inner circle, friends, acquaintances, strangers — and the resulting four-point curve is decomposed into level, linear contraction, curvature, and a residual noise term. The working name for the construct is social proximity conditionality; the person-level statistic is a cross-facet Social Contraction Index.

This paper states the construct, the instrument design, the scoring approach, and — before any data exists — the conditions under which I will consider the construct falsified. Two prior findings constrain the claim severely and are addressed directly rather than deferred: Lenhausen, Bleidorn and Hopwood (2023) found essentially no effect (d < .05) of "people in general" versus "close others" reference-group instructions, and Clifton (2014) has already demonstrated a closeness gradient in Big Five expression across social-network members. The contribution claimed here is narrower than the phenomenon: facet-level resolution, signed directional gradient profiles, and closeness as the ordering axis.

Keywords: Big Five, IPIP, personality assessment, frame of reference, trait expression, social proximity, if–then signatures, self-concept differentiation


1. Introduction

The Five-Factor Model is the most successful descriptive framework in personality psychology, and its public-domain operationalizations — chiefly the International Personality Item Pool (Goldberg, 1999; Goldberg et al., 2006) and Johnson's (2014) 120-item inventory — have made large-scale personality measurement effectively free. That success has a structural consequence worth examining: essentially every widely used instrument reports a facet as a single number.

A single number is a claim. It asserts that the respondent has a characteristic level of the trait, and that variation around that level is either small or uninteresting. For most facets and most people that claim is defensible. This paper concerns the cases where it is not, and where the deviation is not noise but the most descriptively important feature of the trait in that person.

1.1 The observation

The construct originated in an ordinary act of test-taking. Answering the IPIP-NEO-120, I found the Altruism and Sympathy items unanswerable in a specific way. The items are written about people in general:

"I am concerned about others." "I sympathize with the homeless." "I feel sympathy for those who are worse off than myself."

My response to these is genuinely low, and my response to the same behaviors directed at anyone inside a defined personal circle is close to maximal. Neither number is a distortion of the other; there is no intermediate value that describes both. Whichever I chose would be a fabrication of a central tendency I do not have.

The resulting score — 33rd percentile Altruism — is technically defensible and substantively misleading. It suggests moderate, uniform helpfulness. The actual pattern is closer to a step function.

1.2 Why this is not measurement error

Two respondents can score identically on Altruism with entirely different behavioral profiles:

These are not the same person. They will behave differently in most situations that matter — in caregiving, in crisis, in relationship formation, in how they allocate finite attention. A measurement model that cannot distinguish them is not merely imprecise; it is describing a person who does not exist.

The differential is not error variance around a true score. It is a structural property of how the trait is organized in that individual, and current instruments systematically discard it.


2. The measurement problem, stated precisely

Let $T_{ij}$ denote respondent i's expression of trait j. Conventional instruments estimate a single parameter $\hat{T}_{ij}$ and treat within-person variation as measurement error plus situational noise.

Suppose instead that expression is a function of the closeness of the target, $c$:

$$T_{ij}(c) = \alpha_{ij} + \beta_{ij} \cdot f(c) + \varepsilon$$

A conventional instrument, whose items reference an unspecified or distant target, recovers something close to $\alpha_{ij} + \beta_{ij} \cdot f(\bar{c})$ for whatever implicit $\bar{c}$ the item wording evokes. It cannot recover $\beta_{ij}$, and it cannot distinguish a respondent with $\beta = 0$ from one with a large $\beta$ whose weighted average happens to land in the same place.

The claim of this paper is that $\beta$ is (a) estimable from self-report, (b) meaningfully variable between persons, and (c) not redundant with existing measures. Each is testable, and §8 states what would count as a failure.

2.1 A note on the implicit frame

It is tempting to describe conventional items as "unframed." Schulze et al. (2021) argue convincingly that they are not: ostensibly context-free personality items carry hidden, unmeasured framings. The IPIP Altruism items are not neutral with respect to target — they are distant-framed, and their distance is simply undeclared. Part of what this design does is make an existing implicit frame explicit and then vary it.


3. Prior art, and the boundaries of the claim

This section is deliberately adversarial toward the proposal. The literature contains several strands close enough that novelty must be claimed narrowly or not at all.

3.1 The most direct threat

Lenhausen, Bleidorn and Hopwood (2023) administered the Mini-IPIP to 1,194 participants seven times under different instruction sets, including people in general and close others as reference groups. Between-person reference frames produced effects of d < .05 — essentially nothing. Only within-person frames (past self, ideal self) produced substantial shifts, up to d = .98.

Taken at face value this result predicts that the present design will find nothing.

The distinction on which this proposal depends is that Lenhausen et al. manipulated the comparison standardrate yourself relative to close others — whereas this design manipulates the behavioral targethow you actually behave toward close others. These are different psychological operations. The first asks the respondent to re-anchor a self-assessment; the second asks about a different set of behaviors.

Frame-of-reference research that manipulates target or domain does move scores substantially. Shaffer and Postlethwaite's (2012) meta-analysis found contextualized personality measures achieved mean criterion validity of .24 against .11 for non-contextualized measures. Lievens, De Corte and Schollaert (2008) identified the mechanism: contextualization raises validity by reducing within-person inconsistency, because a specified frame makes respondents converge on a single referent.

This distinction is carried entirely by item wording, which converts a theoretical defense into an engineering constraint. It is stated as authoring rule 1 in §5.4, it is a pilot checkpoint, and if the present design nonetheless yields effects of d < .05, the honest conclusion is that Lenhausen et al. were measuring the same thing and the construct does not exist.

3.2 The closest positive precedent

Clifton (2014) asked participants to construct a 30-person social network and rate their own Big Five as experienced with each individual member. Contextual ratings showed incremental validity over global self-report in predicting each informant's perception, and — critically for the present proposal — variability was predicted by the other person's network centrality: participants reported being more extraverted, more neurotic, and less conscientious with more central (closer) network members.

A closeness gradient in Big Five expression has therefore already been demonstrated. No claim of novelty for the phenomenon is available, and none is made. Clifton's finding is also the source of a specific prediction adopted here: that Neuroticism facets may gradient in the opposite direction to prosocial facets (§7, H1c).

3.3 Instruments with the same structure

The design of repeating one item set across several relational targets is well established.

Fraley et al.'s (2011) ECR-RS administers 9 items across 4 targets (mother, father, partner, best friend) and is the methodological template adopted here. Its most important lesson is negative: it models general and target-specific factors simultaneously rather than computing differences between targets.

Robinson (2009) found significant Big Five variation across parent, friend and colleague contexts, with Conscientiousness most stable and Extraversion least. Donahue, Robins, Roberts and John (1993) established self-concept differentiation (SCD), the closest existing measured construct: rate the same attributes across roles, then compute the proportion of variance not explained by the first principal component of the role correlation matrix. High SCD predicts poorer adjustment.

SCD is the most important rival, and it differs in a way that is central to this proposal: SCD discards direction by construction. It measures how much a person varies across roles, not which way. The present design retains sign, and sign is precisely what distinguishes a person who is warmer with intimates from one who is colder with them.

SCD is also known to be confounded with mean-level extremity, profile normativeness, neuroticism, and the number of roles rated (see the critique in Current Psychology, 2014). Any gradient index inherits some of these, which is why neuroticism, self-esteem and authenticity appear as discriminant checks in §8.

3.4 Allocation by closeness

Two literatures already model a psychological quantity as a decreasing function of closeness.

Moral expansiveness (Crimston, Bain, Hornsey & Bastian, 2016) has respondents sort 30 entities into four concentric bands of moral concern, scored 3/2/1/0. Its short form MESx (Crimston et al., 2018) uses 10 entities. It measures the allocation of one quantity — moral standing — by distance. The present proposal generalizes the allocation logic to 30 trait facets, which means MESx combined with global Big Five scores is the obvious rival model and must be beaten on incremental validity. MESx is therefore administered.

Social discounting (Jones & Rachlin, 2006) is the closer quantitative analogue. Participants rank their 100 closest people and report how much money they would forgo to benefit the person at rank N. The amount declines hyperbolically, $v = V_0 / (1 + kN)$, with k as a person-level discount parameter.

The lesson taken from this work is to fit a parameter across several ordered distances rather than subtract two points. The lesson rejected is k itself as a headline statistic, for reasons given in §6.3.

Construal Level Theory (Trope & Liberman, 2010) is a threat rather than a support. Greater psychological distance produces more abstract construal, so items about "people in general" will be answered from semantic rather than episodic memory. Some portion of any observed gradient will be construal artifact rather than trait gating. The instrument's third ring exists to bound this (§5.2).

3.5 Theoretical home

The construct is best understood not as a new trait but as a self-reported if–then signature in the sense of Mischel and Shoda's (1995) Cognitive-Affective Personality System. Their behavioral signatures are stable if(situation)→then(behavior) contingencies, and in the canonical Wediko camp study (Shoda, Mischel & Wright, 1994) the situations were defined substantially by who the other person was. That is structurally the same move made here.

Fleeson's (2001) density-distribution account and Whole Trait Theory (Fleeson & Jayawickreme, 2015) provide the licence for treating traits as distributions rather than points, and report within-person state variability at least as large as between-person variability.

Both frameworks come with a caution. Their parameters are derived from intensive repeated sampling of momentary states. A retrospective questionnaire yields a different estimand: a global recollective judgment, subject to semantic trait inference rather than episodic aggregation. If this instrument is claimed to proxy CAPS signatures or Whole Trait contingency, that claim requires an experience-sampling validation study, which has not been performed. It is identified in §9 as the single most valuable follow-up available.

3.6 What is claimed

Given the above, the claim is confined to:

  1. Facet-level resolution. Prior target-conditioned work is domain-level or adjective-level. A facet-resolution map of trait expression by target closeness does not exist.
  2. Signed, directional gradient profiles, where SCD and related dispersion indices discard sign.
  3. Closeness as the ordering axis rather than social role, which connects trait expression to the discounting, moral-circle and construal-level literatures rather than to role theory.

4. The construct

Social proximity conditionality (SPC) is the degree to which a person's expression of a trait depends on the closeness of the target, together with the shape of that dependence.

Three shapes are distinguished, and the distinction is the substantive content of the construct:

Shape Description Interpretation
Flat Expression is approximately constant across targets Dispositional. The conventional model is correct for this person and facet
Smooth decay Expression declines gradually with distance Graded modulation. Normative for most people on most facets
Cliff Expression drops sharply between two adjacent rings A step function. A categorical boundary rather than a gradient

The cliff case is the one that motivated the work and the one conventional instruments most badly misrepresent, because averaging across a step produces a value the respondent never exhibits.


5. Method

5.1 Big Five backbone

The IPIP-NEO-120 (Johnson, 2014) is administered unmodified: 5 domains × 6 facets × 4 items, 5-point Likert. Items are public domain (Goldberg, 1999; Goldberg et al., 2006).

Leaving the backbone untouched is not a convenience. It is what makes the Big Five scores comparable to published norms and to Johnson's public repository of 619,150 item-level cases, and any edit forfeits that comparability.

Nomenclature: "NEO", "NEO-PI-R" and "NEO PI-3" are trademarks of PAR, Inc. "IPIP-NEO-120" is used here only as a citation of a source instrument.

5.2 The rings

Four ordered targets, with fixed respondent-facing wording:

Ring Anchor
0 the people closest to you — the few you would call first
1 people you would call friends
2 people you see most weeks but would not call friends
3 people in general, including people you have never met

Ring 2 exists to bound the construal confound, not to add resolution. It is concrete and distant simultaneously. If an observed gradient were purely a construal-level artifact — abstract targets answered from semantic memory — ring 2 should pattern with ring 3. If it patterns with closeness instead, the gradient is not merely an abstraction effect.

Ring ordering is verified rather than assumed. A single-item Inclusion of Other in the Self measure (Aron, Aron & Smollan, 1992) is administered per ring as a manipulation check. If IOS does not order the rings as intended, the gradient is uninterpretable and this is reported.

5.3 Gated facets

Ten facets, selected to make the central hypothesis falsifiable rather than merely confirmable.

Facet Domain Predicted gradient Role
Trust A steep positive relational
Altruism A steep positive relational
Sympathy A steep positive relational
Friendliness E steep positive relational
Gregariousness E moderate positive relational
Self-Consciousness N no prediction open question
Morality A flat conduct control
Cooperation A flat conduct control
Dutifulness C flat conduct control
Anger N reversed reversal control

Two categories of control carry most of the design's inferential weight.

Conduct controls (Morality, Cooperation, Dutifulness) are hypothesized to be flat because they operate as rules applied uniformly rather than as affective responses that vary by target. Without facets predicted to be flat, H1 is unfalsifiable.

The reversal control (Anger) is the strongest available defense against response-style artifacts. Anger is predicted to be expressed more toward close others — one is short with family and civil with strangers — consistent with Clifton's (2014) finding of higher self-reported neuroticism with more central network members. A uniform response bias cannot produce opposite-signed slopes. If every gated facet gradients in the same direction, gating is not distinguishable from acquiescence or social desirability; an opposite sign on a theoretically motivated facet is.

Self-Consciousness is included with no directional prediction. Whether shame is more acute before strangers or before people whose judgment one cannot escape is genuinely open, and either answer is informative.

5.4 Item authoring rules

These are constraints on the instrument, not stylistic preferences. Violations degrade the construct without producing any visible symptom.

  1. Behavioral target, never comparison standard. "I go out of my way to help people closest to me" — never "compared to people close to me…". This is the distinction on which §3.1 turns.
  2. Strict parallelism across rings. Ring variants of a stem differ only in the target. Any drift in intensity, specificity or social desirability contaminates the gradient. Each stem belongs to a stem family, and each family appears exactly once per ring.
  3. Original IPIP items are never differenced against ring-framed items. Most IPIP items are target-neutral, not target-distant; differencing a neutral item against a close-framed item confounds specificity with closeness. Both ring framings are authored fresh.
  4. Seven-point scale for gradient items (the backbone remains 5-point). Agreeableness facets sit near ceiling toward close others — in the MES, family and friends occupy the maximum band — and truncation at the ceiling manufactures apparent gradients. Per-ring distributions are inspected in piloting.
  5. Ring order is counterbalanced, the randomization seed is stored, and order × ring interactions are tested and reported.
  6. Balanced keying within each gated facet where the content permits, so that acquiescence remains estimable.

5.5 Forms

A full form of approximately 300 responses (backbone 120; gradient module 120 as 10 facets × 3 stems × 4 rings; auxiliary blocks; MESx; IOS; attention checks) and a short form of approximately 120 responses.

The short form is a strict subset: identical item identifiers and versions. This yields a single poolable dataset rather than two, and permits an in-place upgrade in which a respondent who completes the short form answers only the remaining items.

The short form reports domain scores only by default. Two-item facets have reliabilities in the region of .55–.70, which yields 80% confidence bands spanning roughly 50 percentile points; reporting such a value as a facet percentile would misrepresent its precision.

5.6 Auxiliary measures

5.7 Presentation randomization

Two formats are randomly assigned per session and recorded as a covariate.

Grid: one stem with four ring columns; the stem is read once. Separated: each ring-framed item on its own screen, distributed through the assessment.

The grid permits within-row comparison, which may either inflate consistency or exaggerate contrast; neither risk is obviously larger. Randomization converts an unresolvable design argument into a measurable effect.

The format model must include a slope term, not only an intercept, because a grid layout may steepen the very quantity being measured. If the two arms prove non-invariant — differing loadings rather than differing means — no offset can reconcile them and the arms cannot be pooled for gradient comparisons. Configural → metric → scalar invariance testing across formats therefore precedes any pooled analysis.

5.8 Sample and procedure

Self-selected internet volunteers, 18 and over, recruited without incentive. Participation is pseudonymous: an alias and a retrieval code, with no name, email or password required.

Consent is granular and separately recorded for participation, aggregate inclusion, open-dataset inclusion, re-contact, and sensitive-block inclusion; it is confirmed a second time on completion, and non-consenting records are excluded from all releases.

This sample is not representative. Internet volunteer samples for personality feedback skew young, female, and Anglophone. This limits population inference and is disclosed wherever a percentile appears.


6. Scoring and analysis plan

6.1 Backbone scoring

Facet score is the sum of its four items after reverse keying (range 4–20). Domain score is the mean of its six facet sums (range 4–20), following Kajonius and Johnson (2019) — not the sum of 24 items.

Norms are constructed by scoring Johnson's public 619,150-case dataset with the same engine used for live responses, producing cells by sex × age band. This doubles as the primary regression test of the scoring implementation: domain alphas must reproduce Johnson (2014) within ±.01, and the facet alpha distribution must reproduce Kajonius and Johnson (2019) — mean .78, exactly four facets below .70.

Every reported percentile is labeled with its norm cell and its N.

6.2 Why difference scores are not used

The intuitive statistic — inner-circle score minus stranger score — is not viable. For $D = X - Y$ with comparable variances and reliabilities,

$$\rho_{DD} = \frac{\rho - \rho_{XY}}{1 - \rho_{XY}}$$

IPIP-NEO facets have α ≈ .60–.83, and target-framed trait scores remain strongly correlated across frames (Roberts & Donahue, 1994; Robinson, 2009). At α = .78 and $\rho_{XY}$ = .70 this yields ρ_DD = .27; at α = .65 it is negative. Cronbach and Furby's (1970) advice — that questions framed in terms of gain scores are better reframed — applies directly, as do Edwards' (2002) objections regarding ambiguity, confounded effects, and untested constraints.

6.3 Why the hyperbolic parameter is not the headline

Fitting $v = V_0/(1 + kN)$ after Jones and Rachlin (2006) is superficially attractive. It is rejected as the primary statistic for a specific reason: their N is a genuine ratio-scaled rank position among 100 named people, which is what makes the functional form's shape claim meaningful. The four rings here are ordinal categories. Coding them 1,2,3,4 versus 1,5,20,100 changes k by two orders of magnitude, and a parameter whose value is determined by an arbitrary coding decision cannot carry the argument.

Additionally, k is unidentified for flat curves, undefined for non-monotone ones, and severely right-skewed near zero. Hyperbolic and exponential fits are computed and reported as diagnostics, with the ring-to-distance mapping declared alongside them.

6.4 Orthogonal polynomial contrasts

Four equally spaced ordered points admit an exact saturated orthogonal decomposition:

Component Weights Interpretation
Level (1,1,1,1)/4 Mean expression across targets
Linear (−3,−1,1,3) Contraction — the construct
Quadratic (1,−1,−1,1) Curvature: accelerating versus decelerating
Cubic (−1,3,−3,1) Unpredicted by theory — used as a noise probe

The decomposition is exact, closed-form, always defined, and makes no assumption about the metric underlying the ring ordering beyond ordinality. The OLS slope is recovered as $L/10$.

Because no theory predicts a cubic component in a four-point gradient, the cubic contrast supplies a per-person error estimate at no item cost: $\hat{\sigma}^2 = \frac{1}{F}\sum_f C_f^2/20$, pooled across facets. This assumes zero true cubic variance; if a genuine plateau-then-drop shape exists the estimate is conservative, and the assumption is checkable in aggregate once N is large.

Relativized area under the curve is reported for comparability with the discounting literature but is not led with, because unrelativized AUC under equal spacing is a weighted mean and correlates ≈ .99 with level. The linear contrast is orthogonal to level by construction.

6.5 The Social Contraction Index

Contrast reliability is computable in closed form: $\text{rel}(c'Y) = 1 - (c'\Theta c)/(c'\Sigma c)$. Under plausible parameters the per-facet linear contrast reliability is approximately .35 — better than a difference score's .27, and still insufficient for ranking individuals. This is stated plainly rather than presented as a solution.

Aggregation across facets is what makes the measure usable. If standardized per-facet slopes intercorrelate at $\bar{r} \approx .25$, a nine-facet composite reaches $\alpha \approx .75$ by Spearman-Brown.

This assumption is untested. If $\bar{r}$ is nearer .10 the composite reliability is ≈ .50 and individual-level gradient feedback is not defensible. A test–retest study (n ≥ 150, 2–3 week interval) therefore gates any individual-level reporting of gradient results; population-level findings are unaffected.

Per-facet slopes reported to individuals are shrunk toward the population mean by their own reliability, so that a facet measured at rel = .35 is visibly pulled two-thirds of the way toward average.

6.6 Shape classification and uncertainty

Classification proceeds through three gates: an omnibus χ²(3) test against the null of no gradient; a monotonicity check; and a cliff ratio $CR = d_{max}/D$, where a perfectly linear decay yields CR = 1/3 and a perfect step yields CR = 1. Thresholds are calibrated by simulation against known generating processes rather than asserted, and the resulting confusion matrix is published.

A parametric bootstrap over the ring scores supplies per-person uncertainty for the slope, the shape classification, and the cliff location.

A shape label is reported only when bootstrap confidence in the assigned class is ≥ 0.60. Below that threshold the profile is shown with its error bands and described as not distinguishable from noise. A four-point curve will always appear to have a shape; without this rule the instrument would generate confident descriptions of measurement error.

6.7 Quality control

Long-string runs, intra-individual response variability, per-item latency floors, psychometric antonym/synonym consistency, embedded attention checks, acquiescence over balanced pairs, extremity, and an evaluative-bias index are computed for every session. Exclusion rules are pre-registered. The full flag vector — not only a summary tier — is retained and released, so that other researchers may apply their own criteria rather than inherit ours.


7. Hypotheses

H1 — Facets differ systematically in gradient magnitude. H1a. Affective prosocial facets (Trust, Altruism, Sympathy, Friendliness) show substantially non-zero linear contraction. H1b. Conduct facets (Morality, Cooperation, Dutifulness) show contraction near zero, significantly smaller than H1a facets. H1c. Anger shows contraction of opposite sign. H1b and H1c are the falsification-bearing components. H1a alone would be consistent with response bias.

H2 — Gradient magnitude is a stable individual difference. The Social Contraction Index shows test–retest stability adequate for individual-level interpretation over a 2–3 week interval.

H3 — Gradient magnitude has developmental correlates. Steeper contraction on affective facets is associated with self-reported early social environment and attachment-relevant history. Exploratory; sensitive items are optional and separately consented.

H4 — Gradient shape carries predictive information absent from the mean. Contraction predicts outcomes — circle size, caregiving load, social energy expenditure — beyond facet level.


8. Pre-registered falsification conditions

The construct will be considered not supported if any of the following holds. These are recorded before data collection so that they cannot be revised afterward.

  1. The frames do not separate. Ring-framed scores differ by d < .05, replicating Lenhausen et al. (2023). This would indicate the target manipulation is inert.
  2. No differentiation across facets. Conduct controls show contraction comparable to affective facets, and Anger does not reverse. This would indicate a response-style artifact rather than gating.
  3. Measurement non-invariance across rings. Absent metric and scalar invariance, between-ring differences reflect item functioning rather than trait level and are uninterpretable.
  4. Inadequate reliability. The Social Contraction Index fails to reach reliability sufficient for individual interpretation, in which case individual-level reporting is withdrawn and only population-level findings are reported.
  5. No incremental validity. Contraction fails to predict outcomes beyond (a) global unframed trait level, (b) self-concept differentiation computed on the same data, (c) neuroticism, self-esteem and authenticity, and (d) MESx together with IOS contraction.

Condition 5(d) is the sharpest. If gradient contraction adds nothing beyond IOS closeness ratings and moral expansiveness, then the construct reduces to closeness perception plus moral-circle breadth — both already measurable at a fraction of the item cost — and should be reported as such.

Negative results will be published in the same place, and with the same prominence, as positive ones.


9. Limitations

Self-report of cross-situational contingency. The field's position since Mischel is that people report their own contingencies poorly. Convergence with informant reports or experience sampling has not been demonstrated. This is the most serious limitation, and an ESM validation study correlating the questionnaire gradient with experience-sampled contingency is the single most valuable follow-up available.

Construal-level contamination. Distant targets are construed more abstractly. Ring 2 bounds this but does not eliminate it.

Ceiling effects. Prosocial facets sit near ceiling toward close others. The 7-point scale mitigates truncation; residual restriction is examined per ring.

Differential social desirability. Admitting low concern for strangers is more socially permissible than admitting low concern for family. An apparent gradient may partly be a desirability gradient.

Confounding with adjustment. Every cross-target dispersion index in this literature is entangled with neuroticism and maladjustment. Partial associations are reported.

Norm transportability. Johnson's percentiles apply to the unframed instrument. Framed administrations require their own local norms, and no external norms exist for any novel scale. No percentile is reported for a novel scale below n = 300.

Sample. Self-selected internet volunteers, not a probability sample.


10. Data, licensing, and ethics

Data availability. Aggregate statistics are published openly with no gate. Item-level data are released to researchers on request, restricted to records whose owners consented to open-dataset inclusion, with ages banded and geography coarsened. Codebooks are generated from the item bank and scoring definitions rather than maintained by hand, so they cannot drift from the data they describe.

Licensing. Novel items, scoring keys and data releases are dedicated to the public domain under CC0 1.0, matching the IPIP pool this instrument extends and keeping the whole instrument single-licence. Documentation and written analysis are CC BY 4.0. Attribution is requested as an academic norm, not enforced as a licence condition on the items.

Ethics. This is independent research, not conducted under institutional review. It is not clinical, not diagnostic, and not suitable for selection or evaluative decisions about any person. Participation is pseudonymous, consent is granular and re-confirmed, withdrawal is self-service and results in deletion, and published aggregates are suppressed below a minimum cell size with complementary suppression applied.


References

Aron, A., Aron, E. N., & Smollan, D. (1992). Inclusion of Other in the Self Scale and the structure of interpersonal closeness. Journal of Personality and Social Psychology, 63(4), 596–612.

Clifton, A. (2014). Variability in personality expression across contexts: A social network approach. Journal of Personality, 82(2), 103–115.

Crimston, D., Bain, P. G., Hornsey, M. J., & Bastian, B. (2016). Moral expansiveness: Examining variability in the extension of the moral world. Journal of Personality and Social Psychology, 111(4), 636–653.

Crimston, D., Hornsey, M. J., Bain, P. G., & Bastian, B. (2018). Moral expansiveness short form: Validity and reliability of the MESx. PLOS ONE, 13(10), e0205373.

Cronbach, L. J., & Furby, L. (1970). How we should measure "change" — or should we? Psychological Bulletin, 74(1), 68–80.

Donahue, E. M., Robins, R. W., Roberts, B. W., & John, O. P. (1993). The divided self: Concurrent and longitudinal effects of psychological adjustment and social roles on self-concept differentiation. Journal of Personality and Social Psychology, 64(5), 834–846.

Edwards, J. R. (2002). Alternatives to difference scores: Polynomial regression analysis and response surface methodology. In F. Drasgow & N. W. Schmitt (Eds.), Advances in measurement and data analysis (pp. 350–400). Jossey-Bass.

Fleeson, W. (2001). Toward a structure- and process-integrated view of personality: Traits as density distributions of states. Journal of Personality and Social Psychology, 80(6), 1011–1027.

Fleeson, W., & Jayawickreme, E. (2015). Whole Trait Theory. Journal of Research in Personality, 56, 82–92.

Fraley, R. C., Heffernan, M. E., Vicary, A. M., & Brumbaugh, C. C. (2011). The Experiences in Close Relationships—Relationship Structures questionnaire: A method for assessing attachment orientations across relationships. Psychological Assessment, 23(3), 615–625.

Goldberg, L. R. (1999). A broad-bandwidth, public domain, personality inventory measuring the lower-level facets of several five-factor models. In I. Mervielde et al. (Eds.), Personality Psychology in Europe, 7, 7–28.

Goldberg, L. R., Johnson, J. A., Eber, H. W., Hogan, R., Ashton, M. C., Cloninger, C. R., & Gough, H. G. (2006). The International Personality Item Pool and the future of public-domain personality measures. Journal of Research in Personality, 40, 84–96.

Johnson, J. A. (2014). Measuring thirty facets of the Five Factor Model with a 120-item public domain inventory: Development of the IPIP-NEO-120. Journal of Research in Personality, 51, 78–89.

Jones, B., & Rachlin, H. (2006). Social discounting. Psychological Science, 17(4), 283–286.

Kajonius, P. J., & Johnson, J. A. (2019). Assessing the structure of the Five Factor Model of Personality (IPIP-NEO-120) in the public domain. Europe's Journal of Psychology, 15(2), 260–275.

Lenhausen, M. R., Bleidorn, W., & Hopwood, C. J. (2023). Effects of reference group instructions on Big Five trait scores. Assessment.

Lievens, F., De Corte, W., & Schollaert, E. (2008). A closer look at the frame-of-reference effect in personality scale scores and validity. Journal of Applied Psychology, 93(2), 268–279.

Mischel, W., & Shoda, Y. (1995). A cognitive-affective system theory of personality: Reconceptualizing situations, dispositions, dynamics, and invariance in personality structure. Psychological Review, 102(2), 246–268.

Roberts, B. W., & Donahue, E. M. (1994). One personality, multiple selves: Integrating personality and social roles. Journal of Personality, 62(2), 199–218.

Robinson, O. C. (2009). On the social malleability of traits: Variability and consistency in Big 5 trait expression across three interpersonal contexts. Journal of Individual Differences, 30(4), 201–208.

Schulze, J., et al. (2021). Hidden framings and hidden asymmetries in the measurement of personality. Journal of Personality.

Shaffer, J. A., & Postlethwaite, B. E. (2012). A matter of context: A meta-analytic investigation of the relative validity of contextualized and noncontextualized personality measures. Personnel Psychology, 65(3), 445–494.

Shoda, Y., Mischel, W., & Wright, J. C. (1994). Intraindividual stability in the organization and patterning of behavior. Journal of Personality and Social Psychology, 67(4), 674–687.

Trope, Y., & Liberman, N. (2010). Construal-level theory of psychological distance. Psychological Review, 117(2), 440–463.


Citation

Ticknor, R. (2026). Social proximity conditionality: Measuring Big Five trait expression as a function of closeness to the target [Preprint]. Version 1.0.