If this project works, the part that outlives it is not anybody's result page. It is the data.
Everything gets released — items, scoring keys, aggregate data — under CC0. No account, no request form, no gate, no email address to hand over first.
The model
Openpsychometrics.org is the standard here and the reason is instructive. Their datasets have been used in dozens of published papers, and it is not because the data is unusually good. It is because it is frictionless: consent declared at the start and confirmed at the end, non-consenters purged from every release, a plain CSV, a text codebook, and no login.
A dataset behind a request form is a dataset that gets used by people who already know you exist. That is a much smaller set than the people who could do something with it.
Why CC0 rather than a licence with attribution
This was a genuinely open question and I went back and forth.
The instrument is a mix. The Big Five items are public domain. The closeness items are mine. Under a share-alike or attribution licence, anybody adopting it would have to segregate them — "these items are free, those need attribution, track which is which" — and that is exactly the friction that stops institutional adoption, where a mixed-licence instrument means a legal review before anybody can pilot anything.
I cannot relicense the public-domain items, so the only route to a single-licence instrument is to put mine in the public domain too.
CC0 does not mean going uncited. Academic citation runs on norms and convenience, not licence terms — nobody omits the item pool because they legally could. And in the US, facts are not copyrightable anyway, so asserting a restrictive licence over a dataset is partly unenforceable theatre.
What gets published, including the parts that hurt
The reliability actually observed for every scale, including the ones that disappoint. The invariance results, whether or not they pass. Whether the pre-registered falsification conditions were met.
And the quality flags ship with the data rather than being applied to it. Publishing only the records I consider clean would embed my judgment into everybody else's analysis. Ship the flags and the documented thresholds and let a researcher apply their own rules — that is what makes a dataset reusable rather than merely convenient.
The unglamorous part
Two things stand between here and any of that mattering.
The first is a DOI. A self-hosted page is not durable; the canonical version of the instrument this project is built on once moved from a retired university server to a free webhost and orphaned years of citations. A DOI is the difference between a claim somebody can cite in 2035 and a dead link.
The second is respondents. Aggregate figures are suppressed below twenty-five people in any cell, and there are eleven respondents as I write this, so the honest state of the public dataset today is: correct, published, and almost entirely empty.
That is the right order to build it in. It is also a strange thing to be proud of.