Science Journaling Club Founded 2024

REVIEW · META-ANALYSIS · MARINE BIOLOGY

Not Every Shell Dissolves: Sorting the Ocean Acidification Response Literature by Who Is Actually Affected

Written jointly by the Science Journaling Club

Peer-edited by the club review board · LaTeX source · Our calculation · Interactive model

Abstract We extracted 23 calcification and growth effect sizes from 16 published ocean acidification experiments, pooled them with a DerSimonian-Laird random-effects model, and asked whether taxonomic group explains why those experiments disagree with each other as violently as they do. The pooled response is a 13.0% reduction in calcification or growth (95% CI −31.0% to +9.6%), which is the number most readers arrive wanting and the number the rest of this article spends its length qualifying. The interval contains zero. Heterogeneity is extreme: I² = 98.8%, τ² = 0.278, Q = 1860 on 22 degrees of freedom, and the 95% prediction interval for the next experiment anybody runs stretches from −72% to +167%. Taxon is the obvious candidate, and the naive fixed-effect subgroup decomposition appears to confirm it in spectacular fashion, returning Qbetween = 754 on 7 degrees of freedom. We then fed that same test simulated data in which every taxon had been assigned identical true effects, and at our observed level of within-taxon disagreement it announced a significant taxon difference in 97.9% of runs where none existed. The correctly specified mixed-effects moderator test gives QM = 3.14 on 7 degrees of freedom, p = 0.87, so taxon explains essentially none of the heterogeneity, and experiments on one organism disagree with each other about as hard as experiments on separate phyla do. Eleven students extracting numbers from archived data files are not a systematic review with two independent screeners, and we say plainly below what that costs the estimate and what it does not. So the result worth reporting is the variance itself, together with its flat refusal to sort along the taxonomic lines that everyone, ourselves included, expected of it.

Twenty-Three Numbers, One Headline

The chemistry is settled and we have no argument with it. Dissolve carbon dioxide into seawater and you get carbonic acid, which sheds hydrogen ions, which convert carbonate to bicarbonate and leave less carbonate for a shell to draw on. Surface ocean pH has already fallen by about 0.1 units since 1800, and the standard high-emissions projection takes it down another 0.3 to 0.4 by 2100 [22]. Every step of that is thermodynamics, and nobody serious disputes any of it.

What should get disputed, and mostly does not, is the sentence that comes next, which in most retellings runs something close to this: and therefore shell-building organisms will struggle to build shells.

That sentence is a claim about biology derived from a claim about chemistry, and the biology does not follow from the chemistry nearly as tidily as the retelling needs it to. Organisms are not lumps of calcium carbonate parked in a beaker. Most of them pump ions across membranes, hold internal pH gradients against the surrounding water, coat the growing skeleton in organic film, and spend a great deal of metabolic effort making certain that the fluid where crystals actually form bears very little resemblance to the seawater outside, which is how a coral can hold the pH of its calcifying space several tenths of a unit above the water its polyps are sitting in.

So we went and looked.

We pulled 23 usable effect sizes out of 16 published experiments, working from the raw data files their authors archived wherever such a file existed, rather than from the percentages quoted in their abstracts. Pooled with a standard random-effects model, the whole collection says that lowering pH to end-of-century levels changes calcification or growth by −13.0%, 95% confidence interval −31.0% to +9.6%. The interval contains zero.

The rest of this article explains why quoting that figure on its own would be close to a lie, and section nine takes our own best result apart. Ninety-nine percent of the difference between these experiments is real rather than sampling noise, and the obvious explanation, that different animals respond differently, does not survive being tested properly.

A mean summarises only when the things averaged into it belong together. These do not, and finding out what sorts them instead turned out to be the whole job.

What an Effect Size Actually Says

Before any pooling, the unit. Every number in this article is a log response ratio, and being precise about what that quantity contains repays the trouble, because the word "effect" does an enormous amount of unearned work in most popular retellings of this literature. A log response ratio contains a ratio of two group means and nothing else: no duration, no energy budget, no account of what the animal gave up in order to go on calcifying, and no record of whether the control group was itself in decent condition when the experiment began.

Take an experiment with a control group held at present-day carbon dioxide, a treatment group held at some elevated level, and a mean calcification rate measured in each. The log response ratio is

$$y \;=\; \ln\!\left(\frac{\bar{X}_{\text{treatment}}}{\bar{X}_{\text{control}}}\right)$$

and its sampling variance, by the delta method, is

$$v \;=\; \frac{s_t^{2}}{n_t\,\bar{X}_t^{2}} \;+\; \frac{s_c^{2}}{n_c\,\bar{X}_c^{2}}$$

where \(s\) is a standard deviation and \(n\) a sample size [18]. A value of \(y = 0\) means the treated animals calcified at exactly the control rate, negative values mean slower, and the size of the number carries no information about whether slower matters to the animal. Back-transformation to a percentage runs through \(100(e^{y} - 1)\), which makes \(y = -0.223\) a 20% reduction and \(y = -0.693\) a halving, and which is the conversion standing behind every percentage quoted anywhere in this article.

Two properties matter for what follows. First, the ratio is scale-free, which is the only reason a coccolithophore measured in picograms of calcite per cell per day can share an axis with a coral colony measured in milligrams per day. Second, the logarithm makes the scale symmetric, so that a doubling and a halving sit the same distance from zero in opposite directions, which matters in a collection like ours where several rows report increases large enough to swamp the decreases if the arithmetic were done in raw percentages. Arithmetic percentages are not symmetric that way, which is why a forest plot drawn in percentages looks squashed on the left and stretched out on the right.

The variance formula is where the honesty lives. Notice that \(v\) shrinks when \(n\) grows and when the measurements are tight, and that it carries no term whatever for whether the experiment was well designed, ran long enough, used animals already stressed by collection, or held its carbonate chemistry steady for more than a week. Meta-analysis weights by precision, and precision is not the same thing as being right.

How We Decided What Counted

Screening rules, written down before extraction started, because rules invented afterwards are not rules at all but a description of whatever you happened to keep.

Rule one: the treatment has to be near-future. We kept contrasts where the elevated treatment sat between roughly 700 and 1500 µatm partial pressure of carbon dioxide, or where pH was lowered by 0.2 to 0.5 units below that study's own control. Studies that only reported responses at 2800 µatm were excluded for that contrast even when we used the same study at a lower level, since the question here is what happens under conditions the ocean is actually projected to reach. Extreme treatments produce effects that are large, real and uninformative.

Rule two: the response has to be calcification or growth. Survival, photosynthesis, gene expression and behaviour are all interesting and all out of scope, and where a study reported both calcification and a growth proxy we took calcification.

Rule three: one row per paper per taxon. Several papers tested four strains, or two sites, or nine species. Inside a taxonomic group we combined those by inverse-variance weighting before anything else happened, so that no single paper gets to vote nine times in the coral row and no laboratory gets a say proportional to how many aquaria it owns. Ries and colleagues tested eighteen species at once [7], and that study contributes five rows to our ledger, one per taxonomic group, rather than eighteen.

Rule four: variance from the source, or flagged. Twenty of the 23 rows carry variances we computed from dispersion the authors actually reported, and three do not, because those papers gave a percentage in prose and nothing else. Those three are marked IMPUTED throughout and assigned the 75th percentile of the observed variances, a deliberate choice to hand them very little weight rather than an attempt to guess well.

Where the raw data was archived, we used the archive rather than the abstract. Eleven of our sixteen studies deposited individual-level measurements with PANGAEA under the Ocean Acidification International Coordination Centre compilation, which means that a reader who disagrees with a row of our ledger can download the same file, run the same grouping, and find out in an afternoon which of us is wrong. No other route available to a student club would have given us that standard of checkability, and nothing else about this project comes close to mattering as much.

The Ledger

Every extracted row, sorted by effect size. Study, then the taxonomic group we assigned it to, then the log response ratio, then that same quantity as a percentage change with its 95% interval, and last a column saying where the number came from.

No commentary in this section. The table is the section.

Study Taxon lnRR Percent change (95% CI) Provenance
Interval excludes zero, negative
Comeau 2018 [15]Coralline alga−1.330−73.6 (−90.4, −26.9)PANGAEA 892655; Sporolithon durum 0.0376→0.0088, Neogoniolithon sp. 0.0128→0.0040 mg cm−2 d−1
Maier 2009 [5]Coral−0.789−54.6 (−68.8, −34.0)Table 2, Skagerrak, Lophelia pertusa; 0.046±0.010→0.020±0.003 and 0.021±0.004→0.010±0.002 %/d, n=8
Schoepf 2013 [14] IMPUTEDCoral−0.755−53.0 (−68.9, −29.0)Results text: Acropora millepora "decreased by 53%" at 741 µatm; no dispersion given
Gazeau 2007 [2]Mollusc−0.751−52.8 (−67.2, −32.0)Table 1, Mytilus edulis; 10 incubations ≤756 µatm (0.339±0.095) vs 9 at 981–1461 µatm (0.160±0.079)
Comeau 2018 [15]Coral−0.478−38.0 (−58.2, −7.9)PANGAEA 892655; Acropora yongei 1.696±0.433 (n=7) → 1.052±0.493 (n=7)
Ries 2009 [7]Mollusc−0.400−33.0 (−46.5, −15.9)PANGAEA 733947, 903 vs 409 µatm; nine bivalve and gastropod species combined
Comeau 2009 [4]Pteropod−0.325−27.8 (−32.2, −23.1)45Ca uptake, Limacina helicina; 0.36±0.027 → 0.26±0.018, n=10 each
Uthicke 2012 [11]Foraminifer−0.230−20.5 (−29.6, −10.2)PANGAEA 831207, Marginopora vertebralis; two experiments, pH 8.1 vs 7.7 and 7.8
Langer 2009 [6]Coccolithophore−0.167−15.4 (−19.0, −11.6)Table 3, four Emiliania huxleyi strains, PIC production, n=3; strain range −0.007 to −0.439
Lischka 2011 [9]Pteropod−0.144−13.4 (−23.0, −2.7)PANGAEA 761910, shell increment / diameter; 350 µatm (n=18) vs 1100 µatm (n=15), temperatures pooled
Interval spans zero
Courtney 2013 [12]Echinoderm−1.408−75.5 (−97.9, +183.6)PANGAEA 824707, Echinometra viridis; buoyant-mass gain 15.9% (n=7) vs 3.9% (n=8), recomputed by us
Riebesell 2000 [1]Coccolithophore−0.452−36.4 (−59.6, +0.2)PANGAEA 728092; our bins 250–450 µatm (n=22) vs 650–850 µatm (n=8), two species pooled
Long 2013 [13] IMPUTEDCrustacean−0.095−9.1 (−39.8, +37.3)Results text: Tanner crab percent calcium 10% lower at pH 7.8; no dispersion given
Ries 2009 [7]Coral−0.062−6.0 (−14.9, +3.7)PANGAEA 733947, Oculina arbuscula; 0.1962 (n=11) vs 0.1843 (n=12) %/d
Chauvin 2011 [10]Coral−0.044−4.3 (−20.0, +14.6)PANGAEA 771294, Acropora muricata, 413 vs 1134 µatm; reef flat −0.179, back reef +0.051
Edmunds 2020 [16]Coral−0.030−3.0 (−30.1, +34.6)PANGAEA 926042, Acropora hyacinthus whole colonies; large 1046→989, small 188→196 mg/d
Findlay 2010 [8] IMPUTEDCrustacean+0.033+3.4 (−31.6, +56.2)PANGAEA 737438, Semibalanus balanoides; tank means only, 4.06→4.35 and 4.22→4.21 mg g−1 d−1
Gazeau 2007 [2]Mollusc+0.037+3.8 (−6.5, +15.2)Table 2, Crassostrea gigas; 7 incubations 698–861 µatm vs 6 at 1063–1258 µatm
Comeau 2018 [15]Coral+0.220+24.6 (−26.6, +111.7)PANGAEA 892655; Pocillopora damicornis 0.700 (n=5) → 0.872 (n=6)
Interval excludes zero, positive
Ries 2009 [7]Crustacean+0.245+27.8 (+5.6, +54.6)PANGAEA 733947; blue crab +0.326, lobster +0.062, shrimp +0.584, combined
Ries 2009 [7]Coralline alga+0.620+85.9 (+23.0, +180.8)PANGAEA 733947, Neogoniolithon sp.; 0.0955 (n=10) vs 0.1775 (n=10) %/d
Iglesias-Rodriguez 2008 [3]Coccolithophore+0.679+97.3 (+91.7, +103.0)PANGAEA 718841; six experiments, 300 vs 750 µatm, n=3 per cell, combined
Ries 2009 [7]Echinoderm+0.872+139.3 (+38.3, +313.9)PANGAEA 733947; Arbacia punctulata +1.652, Eucidaris tribuloides −0.442, combined

Ten rows fall below zero with intervals that exclude it. Four rows sit above zero with intervals that exclude it, nine rows have intervals wide enough to contain it, and the two extremes of the ledger are −73.6% for a coralline alga and +139.3% for an urchin.

Notes From the Table

Session 1 · the abstract lied to us, gently

First thing we learned, and it cost an afternoon: the percentage printed in an abstract is often not the percentage you get out of that same paper's own archived data file. Lischka and colleagues report shell increment falling 12% at 1100 µatm [9]. Download the archive, take the 350 µatm treatment as the control rather than the 230 µatm one, and you get 13.4%; take only the in-situ temperature arm and you get 5%. All three numbers are honest. They answer slightly different questions about the same animals, and an abstract has room for one of them.

Session 2 · the light column

Second thing, and this one stopped the meeting for twenty minutes. In the archived Ries dataset, every row of the 409 µatm control block carries a photoperiod of 00:24 and a photosynthetically active radiation value of zero. Read literally, that says the controls ran in total darkness, which cannot be true, since two of the eighteen species are photosynthetic algae that would have died within the week. An annotation artefact, then, almost certainly a curator filling a blank field. We left the data in and used the control block as intended. The twenty minutes were well spent all the same, because they are the reason we now read metadata columns instead of skipping straight to the numbers.

Session 3 · the one we threw out

Manullang and colleagues archived a large, clean dataset covering two Okinawan corals across four carbon dioxide levels, and we spent two sessions trying and failing to make it usable. We could not work out what the units of the calcification column meant, the values barely moved across treatments while the paper reported a significant decline, and nobody in the room could reconcile the two. So we dropped it. Not because it is wrong, but because we could not defend including a number we did not understand, and a meta-analysis that includes numbers its authors do not understand is worse than a smaller one.

Session 4 · the argument about oysters

Gazeau reports a 10% decline for Pacific oysters at 740 µatm [2]. We recomputed from their own incubation table and got +3.8%. The difference is that they fit a regression across the whole pCO2 range including values above 2500 µatm, where oyster calcification really does collapse, and read the fitted line off at 740. We binned instead, comparing incubations near 700–860 with incubations near 1000–1260. Theirs is a better description of the dose-response curve. Ours is a better answer to the question this meta-analysis asks. We flagged it and moved on. Three of us are still not happy.

The Pooled Number

Now the thing everybody wants, followed immediately by the reason not to want it.

Pooling all 23 rows with the DerSimonian-Laird random-effects estimator [19] gives a mean log response ratio of −0.1398 with a standard error of 0.1180. In percentage terms, calcification and growth fall by 13.0%, with a 95% confidence interval running from −31.0% to +9.6%, and the test against zero gives z = −1.18 and p = 0.24. On its own terms, this literature does not demonstrate an average effect. The 95% prediction interval, which asks not where the mean sits but where the next experiment would land, runs from −72% to +167%, and a reader who insists on taking one number away from this article should take that interval rather than the mean.

Failing to demonstrate an effect is not the same as demonstrating none, and anyone who quotes the previous sentence without this one is misusing it badly. What the analysis does demonstrate is that the spread between experiments swamps the mean, and Cochran's Q comes out at 1860 on 22 degrees of freedom where a single common truth would have put it near 22. Eighty-five times over. The corresponding I², meaning the share of observed variation attributable to real differences between experiments rather than to sampling error inside them, is 98.8%, which leaves about one part in a hundred for the noise a confidence interval knows how to describe [20].

Read Figure 1 from the bottom. The diamond is the pooled estimate, and it is narrow. The individual squares are the experiments, and they scatter across nearly the whole width of the plot; a diamond that tight sitting beneath a scatter that wide is a warning label rather than a summary.

study taxon % Courtney 2013echinoderm −76 Comeau 2018coralline −74 Maier 2009coral −55 Schoepf 2013coral −53 Gazeau 2007mussel −53 Comeau 2018coral −38 Riebesell 2000cocco −36 Ries 2009mollusc −33 Comeau 2009pteropod −28 Uthicke 2012foram −21 Langer 2009cocco −15 Lischka 2011pteropod −13 Long 2013crustacean −9 Ries 2009coral −6 Chauvin 2011coral −4 Edmunds 2020coral −3 Findlay 2010crustacean +3 Gazeau 2007oyster +4 Comeau 2018coral +25 Ries 2009crustacean +28 Ries 2009coralline +86 Iglesias 2008cocco +97 Ries 2009echinoderm +139 pooled, k = 23 −13.0% (−31.0, +9.6) 95% prediction −1.5 −1.0 −0.5 0 +0.5 +1.0 log response ratio · 0 = no change · −0.69 = halved solid = negative open = positive
Figure 1. Every effect size we extracted, ordered from most negative to most positive. Squares are point estimates, lines are 95% confidence intervals, arrowheads mark intervals running off the axis. Solid squares are reductions and open squares are increases, which is the value-contrast convention used throughout this article. The diamond is the pooled random-effects estimate and the dashed bar below it is the 95% prediction interval, meaning the range in which the true effect of the next experiment should fall. The prediction bar is more than five times the width of the diamond. That gap is the entire argument of this piece, drawn once.

One more number belongs in this section, and it is the number that changed our minds about how the rest of the article had to be written. If you pool the same 23 rows with a fixed-effect model, which assumes a single common truth and weights purely by precision, you get +0.247, a 28% increase. The sign flips. It flips because one study, Iglesias-Rodriguez and colleagues on Emiliania huxleyi, reports six tightly replicated experiments with a standard error of 0.015, and under fixed-effect weighting that single paper carries more weight than the other fifteen combined [3]. Random effects fixes this by adding τ² to every variance, which flattens the weights until no one paper can run away with the answer. When the sign of a headline result depends on which of two entirely standard estimators the analyst happens to reach for, the thing being pooled is not one population. Two populations in a bucket.

Sorting by Who Is Actually Affected

If the spread is real, the next question is whether it has structure, and taxonomy is the obvious candidate, since nothing else on offer is half as intuitive. Corals and crabs are not the same animal. Perhaps the literature is not disagreeing at all, and is instead agreeing, separately and quietly, about eight different kinds of organism that were never going to answer the same way in the first place.

Split the 23 rows into eight taxonomic groups and the group means do look orderly: molluscs at −29.2%, corals at −22.8%, pteropods at −21.5%, then a gap, then coccolithophores at +4.2%, echinoderms at +6.0% and crustaceans at +14.3%. The standard subgroup decomposition, which takes total heterogeneity and subtracts whatever remains inside the groups, returns a between-group Q of 754.3 on 7 degrees of freedom against an expectation of 7, with a p-value near 10−158 and an air of finality that took us some weeks to shake off.

We believed that for about a week.

Here is the plate we kept redrawing on the whiteboard while we believed it. Solid silhouettes are groups whose 95% interval excludes zero on the negative side. Mid-tone silhouettes mark groups whose point estimate is negative but whose interval crosses zero. Open outlines mark groups whose point estimate is zero or positive, and the whole device was built to make a verdict legible at a glance, which is exactly the property that later turned out to be the problem with it.

Coral −22.8% −38.1 to −3.6
k = 7
Pteropod −21.5% −34.2 to −6.2
k = 2
Foraminifer −20.5% −29.6 to −10.2
k = 1
Mollusc −29.2% −54.6 to +10.4
k = 3
Coralline alga −25.7% −89.0 to +400
k = 2
Coccolithophore +4.2% −47.9 to +108.5
k = 3
Echinoderm +6.0% −87.6 to +802
k = 2
Crustacean +14.3% −6.7 to +39.9
k = 3

Three groups have intervals that exclude zero on the negative side, and they are corals, pteropods and the single foraminifer we managed to extract from anywhere in the literature. Two have negative point estimates with intervals wide enough to include no effect. Three point the other way. Figure 2 puts a scale on the same result and, more usefully, draws the individual experiments underneath each group mean, so that a reader can see for herself which group means come out of agreement and which are the arithmetic average of a brawl.

Look at the circles under each line before you look at the lines. Corals span three quarters of the plot. So do coralline algae, echinoderms and coccolithophores. The only row where the circles cluster tightly is crustaceans. That picture is the first hint that the tidy story about taxa is in trouble, and section nine is where the tidy story stops being a story and starts being an artefact of the wrong test.

taxon k I² within Mollusc3 92% Coralline alga2 92% Coral7 79% Pteropod2 86% Foraminifer1 n/a Coccolithophore3 99.8% Echinoderm2 68% Crustacean3 23% −86% −63% 0 +172% +639% change in calcification or growth (log scale, % labels) Q between groups = 754.3 on 7 df
Figure 2. Pooled effect by taxonomic group. Squares are group means, thick lines are 95% confidence intervals, and the small circles below each line are the individual experiments in that group. The right-hand column gives the within-group I², the share of remaining variation that is real rather than sampling noise. Crustaceans are the only group where that number is low: three independent experiments on barnacles, crabs, a lobster and a shrimp broadly agree with one another. Coccolithophores are the opposite case and get their own section. Note the axis is logarithmic in the effect and merely labelled in percent, which is why +172% and −63% sit the same distance from zero.

The Coccolithophore Problem

One species. One response variable. Three experiments. Effects of +97%, −15% and −36%. Within-group I² of 99.8%.

Emiliania huxleyi is a single-celled alga that plates itself in calcite discs, blooms across thousands of square kilometres of the North Atlantic, and is probably the most-studied calcifier on the planet. If any taxon anywhere in this table should have a settled answer by now, after three decades of culture work in laboratories on two continents, it is this one.

It does not. Riebesell and colleagues, working in 2000, grew cultures across a carbon dioxide gradient and watched calcite production fall as carbon dioxide rose [1]; eight years later Iglesias-Rodriguez and colleagues ran the same organism across a comparable gradient and watched it roughly double [3]. A year after that, Langer and colleagues grew four different strains of the same species side by side in one laboratory and found that the four strains disagreed with each other: one flat, three declining by 27% to 36% at the top of the gradient [6].

Figure 3 puts all three on a single axis, each study normalised to its own present-day control so that the units cancel and only the shapes remain.

0 0.5 1.0 1.5 2.0 200 300 400 600 800 1000 pCO₂ (µatm, logarithmic) calcite production, relative to control Iglesias 2008, 6 runs Langer 2009, 4 strains Riebesell 2000 same species · same response variable · I² within group = 99.8% horizontal dashed line = no change from that study's own control
Figure 3. Three laboratories, one species, opposite conclusions. Each line is a separate culture experiment, rescaled so that its own near-present-day carbon dioxide level sits at 1.0 on the vertical axis. The heavy solid lines rise. The dashed lines start flat and fall. The medium line wanders and then drops. Nothing here is fraud or incompetence; the studies differ in strain, in whether carbonate chemistry was manipulated by acid addition or by bubbling, in light, in nutrient supply, and in whether cultures were acclimated for tens of generations before measurement. Those differences are biologically meaningful. Averaging across them produces +4.2%, a number that describes none of the three experiments.

The dispute was contested at the time and stayed contested. Part of it turns on method, since adding acid to seawater and bubbling carbon dioxide through it do not arrive at the same carbonate chemistry even at the same pH, because the two treatments move total alkalinity and dissolved inorganic carbon differently, and coccolithophores respond to the whole carbonate system rather than to pH alone. Part of it turns on strain. Langer's result is the awkward one, because it removes the between-laboratory explanation entirely: four strains, one laboratory, one protocol, four different answers [6], which leaves genotype as a source of variation sitting well inside the level everyone has been calling a taxon.

If four clonal strains of one species in one flask room cannot agree, the phrase "the response of coccolithophores to acidification" is not describing a quantity that exists.

The Test That Was Lying to Us

A between-group Q of 754 on 7 degrees of freedom is an extraordinary claim, and one of us asked the obvious question late in the process: does that test actually work?

So we checked it the way you check any instrument. Build fake data where you already know the answer, run the instrument on it, and see whether it gives the answer you know to be true. We generated eight taxonomic groups, gave every one of them exactly the same true effect, drew three experiments per group, and ran the subgroup decomposition. Any significant result is a false positive by construction. A calibrated test at the 5% level produces one about one time in twenty, which is the whole meaning of the number 5%.

With no within-group heterogeneity at all the test behaved, returning 5.3% false positives across 1500 simulated study programmes, which is as close to the nominal 5% as fifteen hundred runs can reasonably be expected to land. Then we added within-group heterogeneity at the level our own data shows, τ² = 0.28, and ran the whole thing again.

The test declared a significant between-taxon difference in 97.9% of simulated study programmes where, by construction, no between-taxon difference existed anywhere in the data at all.

The reason is not subtle once you see it. The quantity Qtotal minus the sum of Qwithin is computed on fixed-effect weights, which assume that everything inside a group agrees with everything else inside that group. When experiments inside a group disagree wildly, that disagreement has nowhere to go in the arithmetic, so it lands in the between-group term and is reported there as though it were a difference between taxa. The test is measuring within-group heterogeneity and calling it a group difference.

The correctly specified alternative is a mixed-effects moderator test. You first estimate the residual between-study variance, meaning the disagreement that remains after group means are allowed to differ, then compare group means using weights \(1/(v_i + \hat{\tau}^2_{\text{res}})\) so that within-group spread is charged to the error term where it belongs. On our data the residual variance is \(\hat{\tau}^2_{\text{res}} = 0.2792\), which is almost the whole of the total variance rather than some small remainder, and the test statistic is

$$Q_{M} \;=\; \sum_{g} W_{g}\left(\hat{\theta}_{g} - \hat{\theta}\right)^{2} \;=\; 3.14 \quad \text{on } 7 \text{ d.f.}, \quad p = 0.87$$

Three point one four, on seven degrees of freedom, where chance alone would have delivered seven; the taxonomic differences that looked overwhelming are, once within-taxon disagreement is charged to the right account, indistinguishable from nothing at all.

FALSE POSITIVES WHEN EVERY TAXON HAS THE SAME TRUE EFFECT 1500 simulated programmes per bar · 8 groups 5% 0 25% 50% 75% 100% τ² = 0.00 5.3% 3.9% τ² = 0.05 61.8% 11.1% τ² = 0.28 97.9% 11.1% (our data) fixed-effect Q between groups mixed-effects Qₘ3 experiments each · nominal level 5% THE SAME TWO TESTS ON OUR 23 ROWS 1 3 10 30 100 300 1000 test statistic on 7 degrees of freedom, logarithmic 14.07 = 5% critical value Qₘ = 3.14 Q = 754 left of the dashed line means no detectable taxon effect
Figure 4. Why we threw away our own best result. In the top panel every simulated taxon was given an identical true effect, so every rejection counts as a false positive and an honest test should sit on the dashed 5% line. The solid bars are the fixed-effect subgroup decomposition, the one that produced our Q of 754. It behaves when experiments inside a group agree and collapses when they do not, and this literature is the second case. The open bars are the mixed-effects moderator test, which stays roughly usable although it is still somewhat generous at three experiments per group, because it has to estimate the residual variance from very few studies. The lower panel puts both statistics from the real data on one logarithmic axis against the 5% critical value of chi-squared on 7 degrees of freedom. One sits far to the right of it. The other sits comfortably left.

We want to be precise about what this result does and does not overturn, since most of the trouble in this field arrives through imprecision at exactly this point.

It does not say corals are fine. The coral rows still average −22.8% with an interval that excludes zero, pteropods still average −21.5%, and those seven coral experiments did not stop having found what they found the moment we ran a different test on the collection they belong to.

What it says is narrower. Those group means are not separated from one another by more than the noise sitting inside each group, and corals disagree with corals nearly as much as corals disagree with crabs. The coral group's own I² is 78.9%, the mollusc group's 92.2%, the coccolithophore group's 99.8%, and with disagreement on that scale inside every box on the plate, drawing lines between the boxes is drawing lines on water.

One version of this result is genuinely humbling. We came in expecting to find that the field had flattened a taxon-structured story into a single number, and we fully intended to be the people who un-flattened it. What the arithmetic says instead is that the taxon-structured story is itself a flattening, imposed on a literature whose disagreements do not respect phylum boundaries at all.

So where does the variance live? Not in taxonomy, or at least not detectably at this sample size. The candidates that remain are not mysterious: experiment duration, whether animals were acclimated or shocked, food supply, the method used to manipulate carbonate chemistry, life stage, and whether the organism lives somewhere that already swings through a wider pH range every day than the experiment imposed on it. Each of those has a literature. Testing them properly would take a few hundred rows and two independent screeners, which is the subject of a later section and not of this one.

The Funnel, and What It Admits

Meta-analysis carries a standard diagnostic for the worry that the published record and the conducted record are not the same record, and it is the funnel plot. Plot each effect against its precision. If everything that was run got published, the cloud should be a symmetric funnel, with precise studies clustered near the pooled value at the top and imprecise ones scattering evenly to both sides below. A gap in one lower corner suggests that small studies finding results in that direction were written up, looked unpublishable to somebody, and never made it into print.

Egger's regression test formalises the worry by regressing each effect on its standard error and asking whether the intercept differs from zero [21], and ours does. The intercept is +0.403 with a standard error of 0.114, giving t = 3.52 on 21 degrees of freedom and p = 0.002. The slope runs negative. Imprecise studies in this collection report systematically more negative effects than the precise ones do, which is the pattern a publication filter would leave behind.

We want to be careful about what that licenses you to conclude. Three things.

Start with power. Egger's test has very little of it below about twenty studies, and it is not reliable when the underlying effects are genuinely heterogeneous, which ours emphatically are. A significant intercept in the presence of I² = 98.8% can be produced by real between-study differences that happen to correlate with study size, with nothing whatever suppressed and no editor having declined anything.

Next, how much of the asymmetry hangs on a single point. Dropping the single most influential row, Iglesias-Rodriguez, flips the intercept from +0.403 to −0.177 while leaving it statistically distinguishable from zero (p = 0.004), so the test stays significant and changes its mind about which way the record is skewed. An asymmetry whose direction reverses when you remove one study is not a finding that anybody should build anything at all on.

The likeliest explanation is duller than suppression. Precision and effect are correlated in this collection for a reason that has nothing to do with what did or did not get published. The most precise rows are cell-culture experiments on single-celled algae, three replicate flasks per treatment, instruments reading to four significant figures; the least precise are whole-animal experiments on urchins and corals, where individual variation is enormous. Those two kinds of experiment measure different biology, and the biology they measure happens to differ in direction. Confounding, then, rather than censorship.

Iglesias 2008 Courtney 2013 −1.5 −1.0 −0.5 0 +0.5 +1.0 log response ratio 0 0.25 0.50 0.75 1.00 1.25 standard error (precise at top) pseudo 95% limits Egger fit Egger intercept +0.403 (SE 0.114), t = 3.52 on 21 df, p = 0.002 drop Iglesias 2008 and the intercept becomes −0.177, p = 0.004
Figure 5. Funnel plot. Each circle is one effect size, positioned horizontally by its estimate and vertically by its standard error, so precise experiments sit at the top. The shaded triangle is where 95% of estimates should fall if every experiment were measuring the same true effect of −0.14. Most of the cloud sits outside it, which is another way of stating I² = 98.8%. The dotted line is Egger's fitted regression, tilted because imprecise studies here report more negative effects. Two labelled points do most of the tilting, in opposite ways, and we would not report this diagnostic as evidence of publication bias without far more studies than we have.

The Strongest Case Against This Article

Now the arguments we would make if we were trying to take this piece apart, made as well as we can make them, because an objection you have not stated properly is an objection you have not answered.

Objection one, and it is the serious one

Your central result is a failure to detect, and you are presenting it as a discovery. QM = 3.14 with p = 0.87 does not mean taxa respond alike; it means 23 effect sizes spread thinly over 8 groups cannot resolve a taxon effect when the residual variance stands at 0.279. Your sample size is talking, not the biology. Run the power calculation for this design and you will find that it could barely detect a taxon difference of the size you yourselves estimated, which makes your headline an absence of evidence dressed as evidence of absence, inside an article whose entire complaint concerns people overreading statistics.

We think this objection is correct, and we have made it easy to check. The interactive companion to this article simulates exactly this design: eight groups, three experiments each, residual variance 0.28, and a true between-taxon spread of 0.48, which is the spread of our own group means. Power comes out at about 24%. Three times in four, a taxon difference of exactly the size we appear to see would go unnoticed by the very test we are relying on. Reaching 80% power on those settings would take roughly 22 experiments per taxon, meaning about 175 in total, against the 23 we have.

So the honest statement is narrower than "taxon does not matter". The honest statement runs: this literature, as we were able to extract it, cannot tell the difference between a world where taxa respond differently and a world where they do not. We will defend that weaker claim rather than the stronger one, and we still think it worth publishing, because the sentence being repeated in public is not "we cannot tell yet". The sentence being repeated is a number.

The reply has a second half. Low power explains why we failed to detect a taxon effect, but it does not explain why QM came out at 3.14 rather than, say, 11, which would still have been non-significant while at least pointing in the direction of a real difference. A statistic of 3.14 against a null expectation of 7 is lower than chance would typically deliver, which is to say that the group means sit closer together, relative to their own uncertainty, than random assignment would produce.

Objection two

Your pooled estimate is meaningless and you should not have computed it. Standard methodological advice, with I² sitting at 98.8%, is that pooling is inappropriate, and you computed the number anyway, put it at the head of your abstract, then spent ten sections instructing readers to ignore it. Having it both ways.

Fair. Our defence is that the pooled number is the thing readers arrive expecting, and that refusing to compute it in an article of this kind would look a great deal like concealment. So we compute it, print its prediction interval directly underneath, and make the gap between the two the entire point of Figure 1. A reader who takes only the abstract away gets −13.0% with an interval containing zero, which is at least honest about how little it knows.

Objection three

Your taxon groups are an artefact of which studies you happened to find. Four of your eight groups rest on one or two experiments. The crustacean row leans on Ries 2009, whose blue crabs and shrimp were grown at 25 degrees with unlimited food, conditions under which almost anything calcifies well. Change the study set and the taxon means change with it, which is presumably the whole reason those means refuse to separate from one another at all.

Also fair, and the leave-one-out analysis is where we test it. Dropping any single row moves the pooled estimate only between −16.6% and −10.3%, so the overall figure is stable, while the taxon means are much less stable and we have not claimed otherwise anywhere above. Notice, though, that this objection strengthens our conclusion rather than weakening it: if the taxon means move around when you swap studies in and out, then the taxon means are not measuring a property of the taxon.

Objection four

Short experiments on adults miss the actual mechanism of harm. Almost everything you pooled is a weeks-long exposure of adult or juvenile organisms. The damage that matters for populations happens at fertilisation, at larval settlement and across generations, and it works through energy budgets rather than through direct dissolution of anything. A crab calcifying faster under acidification may simply be spending energy it needed elsewhere. Your +14.3% could be a cost, misread as a benefit.

This one we cannot answer, and we think it is probably right. Wood, Spicer and Widdicombe demonstrated exactly that pattern in a brittlestar, where calcification and metabolism both rose under acidification and the animal paid for the rise in muscle wastage [17]. Nothing in a log response ratio on calcification rate can detect a trade-off of that kind, which means our numbers describe one measured quantity under one kind of experiment on one life stage, and say nothing whatever about whether crustaceans will be fine.

Objection five

None of this is new and the specialists have known it for years. Kroeker and colleagues pooled 228 studies in 2013 and reported that the differences among taxonomic groups were not a significant source of variation in calcification, which is your conclusion, reached on a far larger evidence base and a decade earlier. You have spent a term rediscovering a subordinate clause.

Correct, and the subordinate clause deserves to be on the record rather than buried.

That 2013 analysis rests on more than fourteen times our evidence base [22]. Its headline figures by taxon look much like ours, with corals, coccolithophores and molluscs showing the largest calcification reductions and no detectable effect for echinoderms or crustaceans. Then comes the sentence we read four times before believing it, reporting that these very differences were not a significant source of variation in the calcification analysis that produced them.

So the result we spent a term arriving at was already in print, stated plainly, in the most-cited meta-analysis in the field, and we did not find it because nobody repeats it. What gets repeated from that paper is the taxon-by-taxon table, which is memorable and visual and fits on a slide, while the sentence saying that the table's rows are not statistically distinguishable sits in a subordinate clause in the results section.

Our contribution, if there is one, is the calibration check in section nine. Kroeker and colleagues reached a non-significant subgroup result and reported it. We reached a wildly significant one using a test that should not have been used, caught it by running that test on data with a known answer, and ended up agreeing with them from the opposite direction. The route matters, because it shows how a real analyst using a standard, widely published procedure can be handed a p-value of 10−158 for an effect that is not there.

What We Are Not, and the Finding Anyway

The limits first, stated flatly, because a club that hides its limits has no business publishing a pooled estimate, and because the limits here are large enough to change how much of the rest you should believe.

We are eleven people with library access and a text editor. A systematic review has two people screening every candidate study independently, a pre-registered protocol, a documented search string run across several databases, and a formal risk-of-bias assessment for each study it includes. We had one search, run by whoever was free that week, conducted mostly through a data repository rather than a literature database, and no independent duplicate screening at any stage of it. No amount of correct arithmetic downstream repairs that, and it is the single largest weakness in the article, larger by some distance than anything in the arithmetic.

What follows from that, concretely. Our 16 studies are not a sample of the acidification literature; they are a sample of the acidification literature that archived usable data in a place we could find it, which biases toward European and North American groups, toward the period after about 2007, and toward organisms that survive in aquaria. Three of our 23 rows have variances we invented. Our grouping decisions, our pCO2 window and our choice of control level were all made by us, in a room, over a few weeks, and several of them could reasonably have gone the other way. The oyster disagreement in section five is a live example: our number and the original authors' number differ by fourteen percentage points because of a binning choice.

So how much weight does −13.0% deserve? Very little as an estimate of anything in the ocean. Rather more as a demonstration that the literature does not support a confident single figure, because a biased sample that still fails to agree with itself is telling you something real about disagreement, whatever it fails to tell you about the sea.

And now the finding.

The variation is not noise to be averaged away, and not an embarrassment to be apologised for in a limitations paragraph; the variation is the signal, and it does not run where everybody assumes it runs. Ninety-nine percent of the differences between these experiments are real rather than sampling error, and almost none of that is explained by which organism happened to be sitting in the tank when the pH came down. Two coral experiments differ about as much as a coral and a crab.

A more uncomfortable result than "some taxa are hit and others are not", and a more useful one, because if taxonomy explained the spread the research programme would simply be to work through the phyla one at a time. Since it does not, the programme has to be about the experiments themselves: how long they ran, what the animals were fed, how the carbonate chemistry was manipulated, what life stage went into the tank, and how the organism's home habitat compares with the treatment it was given. Those axes are testable. No pooled percentage shows any of them, and neither does a taxon-by-taxon table, which is why we would rather leave you with a question than a figure.

Here is the plate one last time, drawn the way the arithmetic actually leaves it, with every group open, because no group is separated from the others by more than the noise inside it, and with each cell carrying its own within-group I² in place of the verdict we spent most of a term wanting to give it.

Coral −22.8% I² = 78.9%
k = 7
Pteropod −21.5% I² = 86.0%
k = 2
Foraminifer −20.5% k = 1, no test
Mollusc −29.2% I² = 92.2%
k = 3
Coralline alga −25.7% I² = 91.8%
k = 2
Coccolithophore +4.2% I² = 99.8%
k = 3
Echinoderm +6.0% I² = 68.4%
k = 2
Crustacean +14.3% I² = 23.0%
k = 3

The title of this piece is a small provocation and we will still defend it, though not quite in the sense we meant when we chose it. Not every shell dissolves. Some do, some thin, some thicken, and one urchin got heavier. Saying so is not scepticism about acidification; it is the ordinary business of reading a literature carefully enough to notice that it contains more than one result, and then carefully enough again to notice that the results do not line up along the axis the story requires them to line up along.

References

  1. Riebesell, U., Zondervan, I., Rost, B., Tortell, P. D., Zeebe, R. E. & Morel, F. M. M. (2000). Reduced calcification of marine plankton in response to increased atmospheric CO2. Nature 407, 364–367. doi:10.1038/35030078. Data: PANGAEA doi:10.1594/PANGAEA.728092
  2. Gazeau, F., Quiblier, C., Jansen, J. M., Gattuso, J.-P., Middelburg, J. J. & Heip, C. H. R. (2007). Impact of elevated CO2 on shellfish calcification. Geophysical Research Letters 34, L07603. doi:10.1029/2006GL028554
  3. Iglesias-Rodriguez, M. D., Halloran, P. R., Rickaby, R. E. M., Hall, I. R., Colmenero-Hidalgo, E., Gittins, J. R., Green, D. R. H., Tyrrell, T., Gibbs, S. J., von Dassow, P., Rehm, E., Armbrust, E. V. & Boessenkool, K. P. (2008). Phytoplankton calcification in a high-CO2 world. Science 320, 336–340. doi:10.1126/science.1154122. Data: PANGAEA doi:10.1594/PANGAEA.718841
  4. Comeau, S., Gorsky, G., Jeffree, R., Teyssié, J.-L. & Gattuso, J.-P. (2009). Impact of ocean acidification on a key Arctic pelagic mollusc (Limacina helicina). Biogeosciences 6, 1877–1882. doi:10.5194/bg-6-1877-2009
  5. Maier, C., Hegeman, J., Weinbauer, M. G. & Gattuso, J.-P. (2009). Calcification of the cold-water coral Lophelia pertusa under ambient and reduced pH. Biogeosciences 6, 1671–1680. doi:10.5194/bg-6-1671-2009
  6. Langer, G., Nehrke, G., Probert, I., Ly, J. & Ziveri, P. (2009). Strain-specific responses of Emiliania huxleyi to changing seawater carbonate chemistry. Biogeosciences 6, 2637–2646. doi:10.5194/bg-6-2637-2009
  7. Ries, J. B., Cohen, A. L. & McCorkle, D. C. (2009). Marine calcifiers exhibit mixed responses to CO2-induced ocean acidification. Geology 37, 1131–1134. doi:10.1130/G30210A.1. Data: PANGAEA doi:10.1594/PANGAEA.733947
  8. Findlay, H. S., Kendall, M. A., Spicer, J. I. & Widdicombe, S. (2010). Relative influences of ocean acidification and temperature on intertidal barnacle post-larvae at the northern edge of their geographic distribution. Estuarine, Coastal and Shelf Science 86, 675–682. doi:10.1016/j.ecss.2009.11.036
  9. Lischka, S., Büdenbender, J., Boxhammer, T. & Riebesell, U. (2011). Impact of ocean acidification and elevated temperatures on early juveniles of the polar shelled pteropod Limacina helicina: mortality, shell degradation, and shell growth. Biogeosciences 8, 919–932. doi:10.5194/bg-8-919-2011
  10. Chauvin, A., Denis, V. & Cuet, P. (2011). Is the response of coral calcification to seawater acidification related to nutrient loading? Coral Reefs 30, 911–923. doi:10.1007/s00338-011-0786-7
  11. Uthicke, S. & Fabricius, K. E. (2012). Productivity gains do not compensate for reduced calcification under near-future ocean acidification in the photosynthetic benthic foraminifer species Marginopora vertebralis. Global Change Biology 18, 2781–2791. doi:10.1111/j.1365-2486.2012.02715.x
  12. Courtney, T., Westfield, I. & Ries, J. B. (2013). CO2-induced ocean acidification impairs calcification in the tropical urchin Echinometra viridis. Journal of Experimental Marine Biology and Ecology 440, 169–175. doi:10.1016/j.jembe.2012.11.013
  13. Long, W. C., Swiney, K. M., Harris, C., Page, H. N. & Foy, R. J. (2013). Effects of ocean acidification on juvenile red king crab (Paralithodes camtschaticus) and Tanner crab (Chionoecetes bairdi) growth, condition, calcification, and survival. PLoS ONE 8, e60959. doi:10.1371/journal.pone.0060959
  14. Schoepf, V., Grottoli, A. G., Warner, M. E., Cai, W.-J., Melman, T. F., Hoadley, K. D., Pettay, D. T., Hu, X., Li, Q., Xu, H., Wang, Y., Matsui, Y. & Baumann, J. H. (2013). Coral energy reserves and calcification in a high-CO2 world at two temperatures. PLoS ONE 8, e75049. doi:10.1371/journal.pone.0075049
  15. Comeau, S., Cornwall, C. E., DeCarlo, T. M., Krieger, E. & McCulloch, M. T. (2018). Similar controls on calcification under ocean acidification across unrelated coral reef taxa. Global Change Biology 24, 4857–4868. doi:10.1111/gcb.14379. Data: PANGAEA doi:10.1594/PANGAEA.892655
  16. Edmunds, P. J. & Burgess, S. C. (2020). Emergent properties of branching morphologies modulate the sensitivity of coral calcification to high PCO2. Journal of Experimental Biology 223, jeb217000. doi:10.1242/jeb.217000. Data: PANGAEA doi:10.1594/PANGAEA.926042
  17. Wood, H. L., Spicer, J. I. & Widdicombe, S. (2008). Ocean acidification may increase calcification rates, but at a cost. Proceedings of the Royal Society B 275, 1767–1773. doi:10.1098/rspb.2008.0343
  18. Hedges, L. V., Gurevitch, J. & Curtis, P. S. (1999). The meta-analysis of response ratios in experimental ecology. Ecology 80, 1150–1156. doi:10.1890/0012-9658(1999)080[1150:TMAORR]2.0.CO;2
  19. DerSimonian, R. & Laird, N. (1986). Meta-analysis in clinical trials. Controlled Clinical Trials 7, 177–188. doi:10.1016/0197-2456(86)90046-2
  20. Higgins, J. P. T. & Thompson, S. G. (2002). Quantifying heterogeneity in a meta-analysis. Statistics in Medicine 21, 1539–1558. doi:10.1002/sim.1186
  21. Egger, M., Davey Smith, G., Schneider, M. & Minder, C. (1997). Bias in meta-analysis detected by a simple, graphical test. BMJ 315, 629–634. doi:10.1136/bmj.315.7109.629
  22. Kroeker, K. J., Kordas, R. L., Crim, R., Hendriks, I. E., Ramajo, L., Singh, G. S., Duarte, C. M. & Gattuso, J.-P. (2013). Impacts of ocean acidification on marine organisms: quantifying sensitivities and interaction with warming. Global Change Biology 19, 1884–1896. doi:10.1111/gcb.12179