INTERACTIVE COMPANION · METHODS · STATISTICS
The Selection Bench
Two benches below. The first lets you cut a population wherever you like and watch what happens to the group you cut out when nothing is done to it. The second hands a real treatment effect to two study designs and reports what each one says.
Both run the same model as the club's analysis code, in your browser, offline, with no libraries. Every number they print can be checked by hand against the formula printed beside it. On the settings they load with, the first bench reproduces the article's headline: an apparent gain of 0.8775 standard deviations from an intervention that does not exist.
Model 1. Cut a population and see what moves
Each dot is one simulated individual. Its horizontal position is a first measurement, its vertical position a second measurement of the same person. Between the two, nothing happens. The underlying trait is frozen; only the measurement noise is redrawn.
Drag the vertical line, or use the threshold slider, to choose how much of the population to select. The selected dots light up, their mean on each occasion is marked, and the shift between the two is printed against the closed form
$$\Delta = (1 - r)\,\frac{\varphi(z_p)}{p}, \qquad r = \frac{1}{1 + \lambda^2}$$where \(p\) is the fraction selected, \(z_p\) is the normal quantile at \(p\), and \(\lambda\) is the ratio of noise to signal in the measurement.
Why the top of the class falls at the same time
The bench above has a button hidden in plain sight: the threshold slider also runs the other way conceptually. Whatever gain the bottom \(p\) shows, the top \(p\) shows the same size of decline, because the arithmetic is symmetric about the mean. Nobody publishes that half. It is the same result.
Model 2. Give it a real treatment and watch one design lie
Now there is a genuine effect. A constant shift \(\tau\), in standard deviations, added to the second measurement of anyone treated. The bench enrols the bottom \(p\) of a screening population and runs two designs on the same enrolled people, over and over:
- Single arm, before and after. Everybody enrolled is treated. The estimate is the mean change from first measurement to second. Its expected value is \(\tau + (1-r)\varphi(z_p)/p\).
- Randomised. The enrolled group is split in half at random, one half treated, and the estimate is the difference between the halves on the second measurement. Its expected value is \(\tau\), because both halves regress by exactly the same amount and the comparison subtracts it away.
Press Run. Each trial adds one estimate to each histogram.
Three things to try
Set λ to zero. Perfect measurement. The scatter collapses onto the diagonal, the gain reads exactly 0.0000 at every threshold, and in the trial bench the two designs agree. Regression to the mean is entirely a consequence of measurement error, and with none of it there is none of the artefact.
Set τ to zero and leave everything else alone. The single-arm design reports 0.88 standard deviations of improvement from a treatment that does nothing, and with 200 enrolled it reaches statistical significance essentially every time. The randomised design reports zero. Both studies had the same participants.
Push the threshold to 2% and the noise to λ = 1.5. The apparent gain goes past 1.9 standard deviations. Selecting hard on a bad instrument is the recipe, and both halves of it are things a study does deliberately, for reasons that sound like care.
What this model leaves out
The trait is frozen between measurements, so there is no natural history, no recovery, no learning and no drift. Noise is normal, the same size on both occasions, and independent between them. The treatment effect is the same for everybody. Nothing here is a real patient or a real student, and the reliability \(r\) is handed to the model rather than estimated from anything, which in practice is the hard part. The article's closing section says where each of those choices would move the answer.