There is a sentence most personality tests will never write, and it took me three attempts to get it right in mine.
The sentence is: we could not tell.
How I got it wrong
The instrument classifies each trait's four readings into a shape. One of the available answers was flat — this trait looks the same whoever you are with.
The way that got decided was a significance test. Run the statistic, and if it fails to find a difference, call it flat. This is such a natural thing to write that I did not think about it for a week.
It is wrong, and it is the specific wrongness this whole project exists to complain about. Failing to detect a difference is not evidence that there is no difference. On a short scale with three statements per level, it is mostly evidence that you used a short scale.
A calibration run put a number on it. About nine in ten curves labelled "flat" were gradients the instrument could not resolve. And the page was telling those people that the usual single-number model described them well — the exact overclaim this instrument was built to expose, delivered with confidence, to the people whose data supported it least.
The fix, and what it cost
Sameness now has to be demonstrated, not left over. The test asks whether the readings are close enough together that a real difference of any meaningful size is ruled out — an equivalence test rather than a failed difference test.
And the honest result is brutal: at three statements per level, 2 flat labels survive out of 31,200 simulated respondents. Sameness is essentially never claimable with this instrument.
Reaching a one-point margin would need about twenty-one statements per level instead of three. That is a different, much longer questionnaire.
So the page now says "not separable from noise" where it used to say "steady across all four groups". That is worse for me — one of my ten traits reading as a confident finding was nicer than three of them reading as a shrug. It is also true, and the previous version was not.
Two words that must not be merged
The instrument now distinguishes:
- consistent — the readings agree, and the measurement was precise enough for that to mean something.
- insufficient — the readings did not visibly differ, and the instrument could not have detected it if they had.
Those look similar on a page and they are opposites. Collapsing them lets a failure to measure be reported as a finding of stability, which is the single easiest way to make an instrument lie — and the reader is looking at a chart that visibly disagrees with the caption.
Anywhere a claim is dangerous enough to need a gate, the question worth asking is which other places make the same claim. I found two surfaces making that one, months apart, and only because somebody read their own results closely enough to notice the page contradicting itself two sections apart.