INTERACTIVE COMPANION · META-ANALYSIS · MARINE BIOLOGY
The Taxon Bench
Twenty-three effect sizes, sixteen published experiments, eight taxonomic groups. The article pools them, finds that the pooled number is nearly useless, then finds that taxonomy does not rescue it either. This page lets you check both claims by taking the pool apart with your hands.
Two benches sit below. The first re-runs the whole meta-analysis live while you switch taxa in and out, so you can watch heterogeneity respond. The second stops using the real data and asks a design question instead: if a between-taxon difference of a given size really exists, how many experiments would you have needed to find it?
Model 1. Take the pool apart
Every row below is one study contributing one effect size for one taxon, with the variance computed from the dispersion that study reported. The effect size is the log response ratio \(y = \ln(\bar{X}_t / \bar{X}_c)\), so zero means no change and \(-0.693\) means halved.
Pooling uses the DerSimonian-Laird random-effects estimator. The between-study variance is
$$\hat{\tau}^{2} \;=\; \max\!\left(0,\; \frac{Q - (k-1)}{\sum w_i - \sum w_i^{2} / \sum w_i}\right), \qquad w_i = 1/v_i$$and the pooled estimate then weights each row by \(1/(v_i + \hat{\tau}^2)\). Turning a taxon off removes its rows before any of that is computed, which is why every statistic in the readout moves at once.
- Defaults reproduce
- k = 23, pooled −13.0%, 95% CI −31.0% to +9.6%
- Heterogeneity
- Q = 1860.0 on 22 df, I² = 98.8%, τ² = 0.2775
- Does taxon explain it
- Qₘ = 3.14 on 7 df, p = 0.87. No.
- The naive test
- Q between groups = 754.3 on 7 df, and it is miscalibrated
- Fixed effect
- +28.0% on the same 23 rows, opposite sign
Four things worth trying, in order. Turn off coccolithophores: I² falls from 98.8% to 86.0%, so that group carries a large share of the disagreement without carrying all of it. Turn off corals, the largest group, and the pooled estimate moves from −13.0% to −7.6% while I² climbs to 99.1%. Watch the Qₘ readout the whole time, because it stays resolutely small no matter which groups you keep. Then switch the estimator, which is the one change that alters the answer's sign.
Why turning a group off barely moves the diamond
The instinct is that dropping seven coral rows out of twenty-three should shift things enormously. It shifts the pooled estimate by about five percentage points, which is less than the width of its own confidence interval, and the reason is in the weights.
Under random effects each row is weighted by \(1/(v_i + \hat{\tau}^2)\). Here \(\hat{\tau}^2 = 0.278\), which is larger than the within-study variance of nineteen of our twenty-three rows. When the added constant dominates the denominator, every weight converges on the same value and the pooled estimate becomes close to a plain unweighted average. A plain average of a wide, roughly balanced scatter is stable against losing a few points, and it is also close to meaningless.
That is a property of the method rather than a fact about oceans. It is also the clearest possible signal that the model being fitted does not match the data being fitted to it.
Model 2. How many experiments would it take?
Now leave the real data behind. Suppose there genuinely is a between-taxon difference of a stated size, and suppose you are designing the study programme that has to find it. How many experiments per taxon do you need before the subgroup test reliably notices?
The simulation builds \(G\) taxonomic groups whose true mean effects are spread evenly across a range of width \(\Delta\). Inside each group it draws \(n\) study-level true effects from a normal distribution with variance \(\tau^2_{\text{within}}\), then draws each observed effect around its own true value with sampling variance \(v\). It then runs a subgroup test and records whether that test rejects homogeneity at \(p < 0.05\). Power is the fraction of simulated study programmes in which it does.
Two tests are available, and the difference between them is the whole point of section nine of the article. The naive one subtracts within-group heterogeneity from total heterogeneity on fixed-effect weights:
$$Q_{\text{between}} \;=\; Q_{\text{total}} - \sum_{g} Q_{\text{within},g}, \qquad \mathrm{df} = G - 1$$The correct one estimates a residual variance \(\hat{\tau}^2_{\text{res}}\) first, then compares group means with weights \(1/(v_i + \hat{\tau}^2_{\text{res}})\) so that within-group spread is charged to the error term:
$$Q_{M} \;=\; \sum_{g} W_{g}\left(\hat{\theta}_{g} - \hat{\theta}\right)^{2}, \qquad W_g = \sum_{i \in g} 1/(v_i + \hat{\tau}^2_{\text{res}})$$Three settings deserve a deliberate visit.
Set Δ to zero and leave the correct test selected. Every group now has an identical true effect, so every rejection is a false alarm, and the curve should sit low. It lands near 12% at three experiments per group rather than the nominal 5%, drifting upward as n falls because the residual variance is itself estimated from very few studies. Honest, imperfect, usable.
Keeping Δ at zero, switch to the naive fixed-effect test. The curve leaps to about 98% and stays there at every sample size. Essentially every study programme in which no taxon difference exists would be reported as showing one. That single control is what made us delete a section of the article.
Last, restore the defaults and drag the within-taxon τ² down to 0.05. The required sample size drops sharply. What that says about the real literature is uncomfortable: the reason acidification experiments struggle to resolve taxonomic differences is not that the differences are small. It is that the experiments disagree with each other so violently inside every taxon that a real signal has to shout to be heard over the disagreement.
What this page cannot tell you
Both models take our extraction at face value. If a row is wrong, the pool is wrong, and nothing you do with these sliders will reveal that. The second model is a simulation of a statistical procedure rather than of an ocean; it assumes normal errors, independent studies and a clean group structure, and real acidification experiments violate all three of those in ways we have not modelled.
What the page is good for is one specific thing: showing that the conclusions in the article are not conjured by a particular analytic choice. Switch the estimator, drop the imputed rows, cut the imprecise studies, remove whole taxa, and the scatter stays wide while Qₘ stays small. That stubbornness is the finding.