An honest close, and the entry I expect to rewrite most often.
The assumption everything rests on
Every reliability figure in the closeness module depends on how strongly the four readings correlate with each other within a person. That quantity has never been measured. It is currently modelled, with adjacent levels correlating most.
The model is reasonable and it is still a model. And the module is sensitive to it: the same formula gives a slope reliability of .73 at one plausible value and .15 at another. Every downstream number — the shrunken slopes, the headline index, the width of every band on a closeness chart — inherits it.
There is a second one like it. The person-level index assumes the per-trait slopes correlate at about .25 with each other. That becomes measurable at roughly three hundred responses. If the real figure is nearer .10, the composite falls to about .50 and stops being reportable for an individual at all — and the site will say so rather than quietly continuing to show it.
The instrument is short
Three statements per level. That is few enough that the instrument cannot rule out differences smaller than about two points on a seven-point scale, which means "no gradient detected" is nearly always uninformative rather than a null result.
The fix is more statements, which means a longer questionnaire, which means fewer people finish it. I do not think that trade is obviously worth making yet, and I notice that "not yet" is what somebody says when they do not want to rewrite an item bank.
Nobody has checked the items
The closeness items have been read by one person, who wrote them. Not a psychometrician, not a pilot panel, nobody.
That is conceded publicly on this site and in the file that tells language models what is here, and it is the cheapest weakness on this list to fix. Ten people telling me the wording is confounded would be worth more right now than two hundred more completions.
The controls did not stay flat
The three traits predicted not to move — the conduct rules, the ones that exist so the hypothesis can fail — drift downward in the pilot. Less than the traits predicted to move, but not by nothing.
If that survives more data, it means part of what this instrument sees is a general "I engage less with people further away" effect rather than something specific to any trait. That would not kill the wider thesis, which is only that one number is lossy. It would substantially complicate the structural claim about which traits gate and why.
It is eleven people. It is also the single most interesting thing in the data so far, and it points the wrong way for me, which is a decent sign it is worth taking seriously.
And the ordinary one
Somewhere in the scoring there is a choice between dividing by n and dividing by n − 1 that I have flagged three times and not yet resolved. If it is wrong, every error bar on every closeness chart is about 18% too narrow.
Too narrow is the one direction this project cannot afford to be wrong in. It is at the top of the list, and it has been at the top of the list for three sessions, which is roughly how these things go.