Science Journaling Club Founded 2024

INTERACTIVE COMPANION · WINTER 2026 · STATISTICS

The Forking Paths Bench

A live model accompanying “How Many Reasonable Choices Does It Take to Find a Result That Is Not There”

← Read the full article

Both benches below run on data with nothing in it. Two groups, drawn from the same normal distribution, differing only by chance. Every significant result either bench reports is false, and you can watch exactly how it was manufactured.

The first bench gives you one study and lets you analyse it. The second runs the analysis you chose across thousands of studies and reports how often it finds something.

Bench 1. One study, analysed by hand

Here is a study. Forty per group at full enrolment, three outcome measures correlated at 0.5, one uninformative covariate, one coin-flip attribute per participant. The planned analysis is a two-sample t test on the first outcome measure at the planned sample size. That is the single tick marked planned in the strip below.

Switch on a researcher degree of freedom and more ticks appear, one for every analysis that choice makes available. The ladder underneath records the smallest p value reachable after each choice you make, in the order you make them, which is the honest record of how a finding gets found.

The outcome correlation and the outlier rule are set by the two sliders in bench 2 further down, since both benches have to be analysing the same kind of study for the comparison between them to mean anything. Leave them alone to stay on the article's settings.

paths available: smallest p: planned analysis p: choices needed:

Two things are worth doing here. First, turn everything off, then switch the freedoms on one at a time and watch the ladder step down. Second, change the order you switch them on in. The bottom of the ladder does not move, but the step that crosses the threshold does, and whichever freedom you happened to use first will look like the guilty one.

Bench 2. The same choices, across thousands of studies

One study proves nothing. The question is what the chosen combination does in the long run, so this bench repeats it. Both benches share the five toggles above: whatever is switched on in bench 1 is what bench 2 measures.

The bar fills from the left as studies come in. Dark green is a study that produced a significant result, and every one of them is a false positive, because there is no effect anywhere in the data. The thin amber mark near the left-hand end of the bar is where the fill would stop if the nominal 5% were being honoured.

studies run: 0 false positive rate: standard error: against nominal 0.05: effective independent tests:
Article headline
0.70662 over 50,000 studies, all five freedoms, r = 0.5, 2 SD
Article baseline
0.05044 over 50,000 studies, no freedoms at all
Nominal paths
180 with everything on
Effective independent tests
23.91, so 7.5 paths per test's worth of damage
If the 180 paths were independent
1 − 0.95180 = 0.99990

This page starts on the article's settings, so the first run should land near 0.707 give or take its standard error. Turn all five off and it should land near 0.05, which is the only result on this page that is supposed to be there.

The arithmetic on the readout

Two of the numbers in the readout deserve explaining, because neither is obvious.

The standard error is binomial. Each study either finds something or does not, so a rate over n studies carries

SE = sqrt( p (1 − p) / n )

which at p = 0.7 and n = 5,000 is 0.0065. Two decimal places is all such a run can support, and the third digit on the readout is there so that you can watch it wobble.

The effective number of independent tests inverts the formula for the chance that at least one of k independent tests at level α comes out significant:

keff = ln(1 − rate) / ln(1 − α)

Set every freedom on and the bench offers 180 analysis paths, but the keff readout will sit near 24. The paths overlap, because they are 180 views of the same 80 numbers, and that overlap is the only thing keeping the rate at 0.71 instead of 0.99990.

What this page cannot show you

Nobody analyses data by enumerating 180 paths and taking the minimum. Both benches do exactly that, so they give an upper bound for these five freedoms rather than a description of anybody's Tuesday afternoon. The honest reading of the number is that it is what the practice is worth if carried to its limit.

In the other direction the benches are far too kind. There is no file drawer here, no choice between a t test and its non-parametric cousin, no log transform, no dropping of a whole condition, and no hypothesis written after the result was known. Published checklists run to more than thirty decision points. This page models five.

And nothing here has a real effect in it anywhere, so the benches say nothing about what these practices do to the size of a finding that is genuinely present. That is a different study, and a harder one.