REVIEW · META-ANALYSIS · MARINE BIOLOGY
Not Every Shell Dissolves: Sorting the Ocean Acidification Response Literature by Who Is Actually Affected
Peer-edited by the club review board · LaTeX source · Our calculation · Interactive model
Twenty-Three Numbers, One Headline
The chemistry is settled and we have no argument with it. Dissolve carbon dioxide into seawater and you get carbonic acid, which sheds hydrogen ions, which convert carbonate to bicarbonate and leave less carbonate for a shell to draw on. Surface ocean pH has already fallen by about 0.1 units since 1800, and the standard high-emissions projection takes it down another 0.3 to 0.4 by 2100 [22]. Every step of that is thermodynamics, and nobody serious disputes any of it.
What should get disputed, and mostly does not, is the sentence that comes next, which in most retellings runs something close to this: and therefore shell-building organisms will struggle to build shells.
That sentence is a claim about biology derived from a claim about chemistry, and the biology does not follow from the chemistry nearly as tidily as the retelling needs it to. Organisms are not lumps of calcium carbonate parked in a beaker. Most of them pump ions across membranes, hold internal pH gradients against the surrounding water, coat the growing skeleton in organic film, and spend a great deal of metabolic effort making certain that the fluid where crystals actually form bears very little resemblance to the seawater outside, which is how a coral can hold the pH of its calcifying space several tenths of a unit above the water its polyps are sitting in.
So we went and looked.
We pulled 23 usable effect sizes out of 16 published experiments, working from the raw data files their authors archived wherever such a file existed, rather than from the percentages quoted in their abstracts. Pooled with a standard random-effects model, the whole collection says that lowering pH to end-of-century levels changes calcification or growth by −13.0%, 95% confidence interval −31.0% to +9.6%. The interval contains zero.
The rest of this article explains why quoting that figure on its own would be close to a lie, and section nine takes our own best result apart. Ninety-nine percent of the difference between these experiments is real rather than sampling noise, and the obvious explanation, that different animals respond differently, does not survive being tested properly.
A mean summarises only when the things averaged into it belong together. These do not, and finding out what sorts them instead turned out to be the whole job.
What an Effect Size Actually Says
Before any pooling, the unit. Every number in this article is a log response ratio, and being precise about what that quantity contains repays the trouble, because the word "effect" does an enormous amount of unearned work in most popular retellings of this literature. A log response ratio contains a ratio of two group means and nothing else: no duration, no energy budget, no account of what the animal gave up in order to go on calcifying, and no record of whether the control group was itself in decent condition when the experiment began.
Take an experiment with a control group held at present-day carbon dioxide, a treatment group held at some elevated level, and a mean calcification rate measured in each. The log response ratio is
$$y \;=\; \ln\!\left(\frac{\bar{X}_{\text{treatment}}}{\bar{X}_{\text{control}}}\right)$$and its sampling variance, by the delta method, is
$$v \;=\; \frac{s_t^{2}}{n_t\,\bar{X}_t^{2}} \;+\; \frac{s_c^{2}}{n_c\,\bar{X}_c^{2}}$$where \(s\) is a standard deviation and \(n\) a sample size [18]. A value of \(y = 0\) means the treated animals calcified at exactly the control rate, negative values mean slower, and the size of the number carries no information about whether slower matters to the animal. Back-transformation to a percentage runs through \(100(e^{y} - 1)\), which makes \(y = -0.223\) a 20% reduction and \(y = -0.693\) a halving, and which is the conversion standing behind every percentage quoted anywhere in this article.
Two properties matter for what follows. First, the ratio is scale-free, which is the only reason a coccolithophore measured in picograms of calcite per cell per day can share an axis with a coral colony measured in milligrams per day. Second, the logarithm makes the scale symmetric, so that a doubling and a halving sit the same distance from zero in opposite directions, which matters in a collection like ours where several rows report increases large enough to swamp the decreases if the arithmetic were done in raw percentages. Arithmetic percentages are not symmetric that way, which is why a forest plot drawn in percentages looks squashed on the left and stretched out on the right.
The variance formula is where the honesty lives. Notice that \(v\) shrinks when \(n\) grows and when the measurements are tight, and that it carries no term whatever for whether the experiment was well designed, ran long enough, used animals already stressed by collection, or held its carbonate chemistry steady for more than a week. Meta-analysis weights by precision, and precision is not the same thing as being right.
How We Decided What Counted
Screening rules, written down before extraction started, because rules invented afterwards are not rules at all but a description of whatever you happened to keep.
Rule one: the treatment has to be near-future. We kept contrasts where the elevated treatment sat between roughly 700 and 1500 µatm partial pressure of carbon dioxide, or where pH was lowered by 0.2 to 0.5 units below that study's own control. Studies that only reported responses at 2800 µatm were excluded for that contrast even when we used the same study at a lower level, since the question here is what happens under conditions the ocean is actually projected to reach. Extreme treatments produce effects that are large, real and uninformative.
Rule two: the response has to be calcification or growth. Survival, photosynthesis, gene expression and behaviour are all interesting and all out of scope, and where a study reported both calcification and a growth proxy we took calcification.
Rule three: one row per paper per taxon. Several papers tested four strains, or two sites, or nine species. Inside a taxonomic group we combined those by inverse-variance weighting before anything else happened, so that no single paper gets to vote nine times in the coral row and no laboratory gets a say proportional to how many aquaria it owns. Ries and colleagues tested eighteen species at once [7], and that study contributes five rows to our ledger, one per taxonomic group, rather than eighteen.
Rule four: variance from the source, or flagged. Twenty of the 23 rows carry variances we computed from dispersion the authors actually reported, and three do not, because those papers gave a percentage in prose and nothing else. Those three are marked IMPUTED throughout and assigned the 75th percentile of the observed variances, a deliberate choice to hand them very little weight rather than an attempt to guess well.
Where the raw data was archived, we used the archive rather than the abstract. Eleven of our sixteen studies deposited individual-level measurements with PANGAEA under the Ocean Acidification International Coordination Centre compilation, which means that a reader who disagrees with a row of our ledger can download the same file, run the same grouping, and find out in an afternoon which of us is wrong. No other route available to a student club would have given us that standard of checkability, and nothing else about this project comes close to mattering as much.
The Ledger
Every extracted row, sorted by effect size. Study, then the taxonomic group we assigned it to, then the log response ratio, then that same quantity as a percentage change with its 95% interval, and last a column saying where the number came from.
No commentary in this section. The table is the section.
| Study | Taxon | lnRR | Percent change (95% CI) | Provenance |
|---|---|---|---|---|
| Interval excludes zero, negative | ||||
| Comeau 2018 [15] | Coralline alga | −1.330 | −73.6 (−90.4, −26.9) | PANGAEA 892655; Sporolithon durum 0.0376→0.0088, Neogoniolithon sp. 0.0128→0.0040 mg cm−2 d−1 |
| Maier 2009 [5] | Coral | −0.789 | −54.6 (−68.8, −34.0) | Table 2, Skagerrak, Lophelia pertusa; 0.046±0.010→0.020±0.003 and 0.021±0.004→0.010±0.002 %/d, n=8 |
| Schoepf 2013 [14] IMPUTED | Coral | −0.755 | −53.0 (−68.9, −29.0) | Results text: Acropora millepora "decreased by 53%" at 741 µatm; no dispersion given |
| Gazeau 2007 [2] | Mollusc | −0.751 | −52.8 (−67.2, −32.0) | Table 1, Mytilus edulis; 10 incubations ≤756 µatm (0.339±0.095) vs 9 at 981–1461 µatm (0.160±0.079) |
| Comeau 2018 [15] | Coral | −0.478 | −38.0 (−58.2, −7.9) | PANGAEA 892655; Acropora yongei 1.696±0.433 (n=7) → 1.052±0.493 (n=7) |
| Ries 2009 [7] | Mollusc | −0.400 | −33.0 (−46.5, −15.9) | PANGAEA 733947, 903 vs 409 µatm; nine bivalve and gastropod species combined |
| Comeau 2009 [4] | Pteropod | −0.325 | −27.8 (−32.2, −23.1) | 45Ca uptake, Limacina helicina; 0.36±0.027 → 0.26±0.018, n=10 each |
| Uthicke 2012 [11] | Foraminifer | −0.230 | −20.5 (−29.6, −10.2) | PANGAEA 831207, Marginopora vertebralis; two experiments, pH 8.1 vs 7.7 and 7.8 |
| Langer 2009 [6] | Coccolithophore | −0.167 | −15.4 (−19.0, −11.6) | Table 3, four Emiliania huxleyi strains, PIC production, n=3; strain range −0.007 to −0.439 |
| Lischka 2011 [9] | Pteropod | −0.144 | −13.4 (−23.0, −2.7) | PANGAEA 761910, shell increment / diameter; 350 µatm (n=18) vs 1100 µatm (n=15), temperatures pooled |
| Interval spans zero | ||||
| Courtney 2013 [12] | Echinoderm | −1.408 | −75.5 (−97.9, +183.6) | PANGAEA 824707, Echinometra viridis; buoyant-mass gain 15.9% (n=7) vs 3.9% (n=8), recomputed by us |
| Riebesell 2000 [1] | Coccolithophore | −0.452 | −36.4 (−59.6, +0.2) | PANGAEA 728092; our bins 250–450 µatm (n=22) vs 650–850 µatm (n=8), two species pooled |
| Long 2013 [13] IMPUTED | Crustacean | −0.095 | −9.1 (−39.8, +37.3) | Results text: Tanner crab percent calcium 10% lower at pH 7.8; no dispersion given |
| Ries 2009 [7] | Coral | −0.062 | −6.0 (−14.9, +3.7) | PANGAEA 733947, Oculina arbuscula; 0.1962 (n=11) vs 0.1843 (n=12) %/d |
| Chauvin 2011 [10] | Coral | −0.044 | −4.3 (−20.0, +14.6) | PANGAEA 771294, Acropora muricata, 413 vs 1134 µatm; reef flat −0.179, back reef +0.051 |
| Edmunds 2020 [16] | Coral | −0.030 | −3.0 (−30.1, +34.6) | PANGAEA 926042, Acropora hyacinthus whole colonies; large 1046→989, small 188→196 mg/d |
| Findlay 2010 [8] IMPUTED | Crustacean | +0.033 | +3.4 (−31.6, +56.2) | PANGAEA 737438, Semibalanus balanoides; tank means only, 4.06→4.35 and 4.22→4.21 mg g−1 d−1 |
| Gazeau 2007 [2] | Mollusc | +0.037 | +3.8 (−6.5, +15.2) | Table 2, Crassostrea gigas; 7 incubations 698–861 µatm vs 6 at 1063–1258 µatm |
| Comeau 2018 [15] | Coral | +0.220 | +24.6 (−26.6, +111.7) | PANGAEA 892655; Pocillopora damicornis 0.700 (n=5) → 0.872 (n=6) |
| Interval excludes zero, positive | ||||
| Ries 2009 [7] | Crustacean | +0.245 | +27.8 (+5.6, +54.6) | PANGAEA 733947; blue crab +0.326, lobster +0.062, shrimp +0.584, combined |
| Ries 2009 [7] | Coralline alga | +0.620 | +85.9 (+23.0, +180.8) | PANGAEA 733947, Neogoniolithon sp.; 0.0955 (n=10) vs 0.1775 (n=10) %/d |
| Iglesias-Rodriguez 2008 [3] | Coccolithophore | +0.679 | +97.3 (+91.7, +103.0) | PANGAEA 718841; six experiments, 300 vs 750 µatm, n=3 per cell, combined |
| Ries 2009 [7] | Echinoderm | +0.872 | +139.3 (+38.3, +313.9) | PANGAEA 733947; Arbacia punctulata +1.652, Eucidaris tribuloides −0.442, combined |
Ten rows fall below zero with intervals that exclude it. Four rows sit above zero with intervals that exclude it, nine rows have intervals wide enough to contain it, and the two extremes of the ledger are −73.6% for a coralline alga and +139.3% for an urchin.
Notes From the Table
Session 1 · the abstract lied to us, gentlyFirst thing we learned, and it cost an afternoon: the percentage printed in an abstract is often not the percentage you get out of that same paper's own archived data file. Lischka and colleagues report shell increment falling 12% at 1100 µatm [9]. Download the archive, take the 350 µatm treatment as the control rather than the 230 µatm one, and you get 13.4%; take only the in-situ temperature arm and you get 5%. All three numbers are honest. They answer slightly different questions about the same animals, and an abstract has room for one of them.
Session 2 · the light columnSecond thing, and this one stopped the meeting for twenty minutes. In the archived Ries dataset, every row of the 409 µatm control block carries a photoperiod of 00:24 and a photosynthetically active radiation value of zero. Read literally, that says the controls ran in total darkness, which cannot be true, since two of the eighteen species are photosynthetic algae that would have died within the week. An annotation artefact, then, almost certainly a curator filling a blank field. We left the data in and used the control block as intended. The twenty minutes were well spent all the same, because they are the reason we now read metadata columns instead of skipping straight to the numbers.
Session 3 · the one we threw outManullang and colleagues archived a large, clean dataset covering two Okinawan corals across four carbon dioxide levels, and we spent two sessions trying and failing to make it usable. We could not work out what the units of the calcification column meant, the values barely moved across treatments while the paper reported a significant decline, and nobody in the room could reconcile the two. So we dropped it. Not because it is wrong, but because we could not defend including a number we did not understand, and a meta-analysis that includes numbers its authors do not understand is worse than a smaller one.
Session 4 · the argument about oystersGazeau reports a 10% decline for Pacific oysters at 740 µatm [2]. We recomputed from their own incubation table and got +3.8%. The difference is that they fit a regression across the whole pCO2 range including values above 2500 µatm, where oyster calcification really does collapse, and read the fitted line off at 740. We binned instead, comparing incubations near 700–860 with incubations near 1000–1260. Theirs is a better description of the dose-response curve. Ours is a better answer to the question this meta-analysis asks. We flagged it and moved on. Three of us are still not happy.
The Pooled Number
Now the thing everybody wants, followed immediately by the reason not to want it.
Pooling all 23 rows with the DerSimonian-Laird random-effects estimator [19] gives a mean log response ratio of −0.1398 with a standard error of 0.1180. In percentage terms, calcification and growth fall by 13.0%, with a 95% confidence interval running from −31.0% to +9.6%, and the test against zero gives z = −1.18 and p = 0.24. On its own terms, this literature does not demonstrate an average effect. The 95% prediction interval, which asks not where the mean sits but where the next experiment would land, runs from −72% to +167%, and a reader who insists on taking one number away from this article should take that interval rather than the mean.
Failing to demonstrate an effect is not the same as demonstrating none, and anyone who quotes the previous sentence without this one is misusing it badly. What the analysis does demonstrate is that the spread between experiments swamps the mean, and Cochran's Q comes out at 1860 on 22 degrees of freedom where a single common truth would have put it near 22. Eighty-five times over. The corresponding I², meaning the share of observed variation attributable to real differences between experiments rather than to sampling error inside them, is 98.8%, which leaves about one part in a hundred for the noise a confidence interval knows how to describe [20].
Read Figure 1 from the bottom. The diamond is the pooled estimate, and it is narrow. The individual squares are the experiments, and they scatter across nearly the whole width of the plot; a diamond that tight sitting beneath a scatter that wide is a warning label rather than a summary.
One more number belongs in this section, and it is the number that changed our minds about how the rest of the article had to be written. If you pool the same 23 rows with a fixed-effect model, which assumes a single common truth and weights purely by precision, you get +0.247, a 28% increase. The sign flips. It flips because one study, Iglesias-Rodriguez and colleagues on Emiliania huxleyi, reports six tightly replicated experiments with a standard error of 0.015, and under fixed-effect weighting that single paper carries more weight than the other fifteen combined [3]. Random effects fixes this by adding τ² to every variance, which flattens the weights until no one paper can run away with the answer. When the sign of a headline result depends on which of two entirely standard estimators the analyst happens to reach for, the thing being pooled is not one population. Two populations in a bucket.
Sorting by Who Is Actually Affected
If the spread is real, the next question is whether it has structure, and taxonomy is the obvious candidate, since nothing else on offer is half as intuitive. Corals and crabs are not the same animal. Perhaps the literature is not disagreeing at all, and is instead agreeing, separately and quietly, about eight different kinds of organism that were never going to answer the same way in the first place.
Split the 23 rows into eight taxonomic groups and the group means do look orderly: molluscs at −29.2%, corals at −22.8%, pteropods at −21.5%, then a gap, then coccolithophores at +4.2%, echinoderms at +6.0% and crustaceans at +14.3%. The standard subgroup decomposition, which takes total heterogeneity and subtracts whatever remains inside the groups, returns a between-group Q of 754.3 on 7 degrees of freedom against an expectation of 7, with a p-value near 10−158 and an air of finality that took us some weeks to shake off.
We believed that for about a week.
Here is the plate we kept redrawing on the whiteboard while we believed it. Solid silhouettes are groups whose 95% interval excludes zero on the negative side. Mid-tone silhouettes mark groups whose point estimate is negative but whose interval crosses zero. Open outlines mark groups whose point estimate is zero or positive, and the whole device was built to make a verdict legible at a glance, which is exactly the property that later turned out to be the problem with it.
k = 7
k = 2
k = 1
k = 3
k = 2
k = 3
k = 2
k = 3
Three groups have intervals that exclude zero on the negative side, and they are corals, pteropods and the single foraminifer we managed to extract from anywhere in the literature. Two have negative point estimates with intervals wide enough to include no effect. Three point the other way. Figure 2 puts a scale on the same result and, more usefully, draws the individual experiments underneath each group mean, so that a reader can see for herself which group means come out of agreement and which are the arithmetic average of a brawl.
Look at the circles under each line before you look at the lines. Corals span three quarters of the plot. So do coralline algae, echinoderms and coccolithophores. The only row where the circles cluster tightly is crustaceans. That picture is the first hint that the tidy story about taxa is in trouble, and section nine is where the tidy story stops being a story and starts being an artefact of the wrong test.
The Coccolithophore Problem
One species. One response variable. Three experiments. Effects of +97%, −15% and −36%. Within-group I² of 99.8%.
Emiliania huxleyi is a single-celled alga that plates itself in calcite discs, blooms across thousands of square kilometres of the North Atlantic, and is probably the most-studied calcifier on the planet. If any taxon anywhere in this table should have a settled answer by now, after three decades of culture work in laboratories on two continents, it is this one.
It does not. Riebesell and colleagues, working in 2000, grew cultures across a carbon dioxide gradient and watched calcite production fall as carbon dioxide rose [1]; eight years later Iglesias-Rodriguez and colleagues ran the same organism across a comparable gradient and watched it roughly double [3]. A year after that, Langer and colleagues grew four different strains of the same species side by side in one laboratory and found that the four strains disagreed with each other: one flat, three declining by 27% to 36% at the top of the gradient [6].
Figure 3 puts all three on a single axis, each study normalised to its own present-day control so that the units cancel and only the shapes remain.
The dispute was contested at the time and stayed contested. Part of it turns on method, since adding acid to seawater and bubbling carbon dioxide through it do not arrive at the same carbonate chemistry even at the same pH, because the two treatments move total alkalinity and dissolved inorganic carbon differently, and coccolithophores respond to the whole carbonate system rather than to pH alone. Part of it turns on strain. Langer's result is the awkward one, because it removes the between-laboratory explanation entirely: four strains, one laboratory, one protocol, four different answers [6], which leaves genotype as a source of variation sitting well inside the level everyone has been calling a taxon.
If four clonal strains of one species in one flask room cannot agree, the phrase "the response of coccolithophores to acidification" is not describing a quantity that exists.
The Test That Was Lying to Us
A between-group Q of 754 on 7 degrees of freedom is an extraordinary claim, and one of us asked the obvious question late in the process: does that test actually work?
So we checked it the way you check any instrument. Build fake data where you already know the answer, run the instrument on it, and see whether it gives the answer you know to be true. We generated eight taxonomic groups, gave every one of them exactly the same true effect, drew three experiments per group, and ran the subgroup decomposition. Any significant result is a false positive by construction. A calibrated test at the 5% level produces one about one time in twenty, which is the whole meaning of the number 5%.
With no within-group heterogeneity at all the test behaved, returning 5.3% false positives across 1500 simulated study programmes, which is as close to the nominal 5% as fifteen hundred runs can reasonably be expected to land. Then we added within-group heterogeneity at the level our own data shows, τ² = 0.28, and ran the whole thing again.
The test declared a significant between-taxon difference in 97.9% of simulated study programmes where, by construction, no between-taxon difference existed anywhere in the data at all.
The reason is not subtle once you see it. The quantity Qtotal minus the sum of Qwithin is computed on fixed-effect weights, which assume that everything inside a group agrees with everything else inside that group. When experiments inside a group disagree wildly, that disagreement has nowhere to go in the arithmetic, so it lands in the between-group term and is reported there as though it were a difference between taxa. The test is measuring within-group heterogeneity and calling it a group difference.
The correctly specified alternative is a mixed-effects moderator test. You first estimate the residual between-study variance, meaning the disagreement that remains after group means are allowed to differ, then compare group means using weights \(1/(v_i + \hat{\tau}^2_{\text{res}})\) so that within-group spread is charged to the error term where it belongs. On our data the residual variance is \(\hat{\tau}^2_{\text{res}} = 0.2792\), which is almost the whole of the total variance rather than some small remainder, and the test statistic is
$$Q_{M} \;=\; \sum_{g} W_{g}\left(\hat{\theta}_{g} - \hat{\theta}\right)^{2} \;=\; 3.14 \quad \text{on } 7 \text{ d.f.}, \quad p = 0.87$$Three point one four, on seven degrees of freedom, where chance alone would have delivered seven; the taxonomic differences that looked overwhelming are, once within-taxon disagreement is charged to the right account, indistinguishable from nothing at all.
We want to be precise about what this result does and does not overturn, since most of the trouble in this field arrives through imprecision at exactly this point.
It does not say corals are fine. The coral rows still average −22.8% with an interval that excludes zero, pteropods still average −21.5%, and those seven coral experiments did not stop having found what they found the moment we ran a different test on the collection they belong to.
What it says is narrower. Those group means are not separated from one another by more than the noise sitting inside each group, and corals disagree with corals nearly as much as corals disagree with crabs. The coral group's own I² is 78.9%, the mollusc group's 92.2%, the coccolithophore group's 99.8%, and with disagreement on that scale inside every box on the plate, drawing lines between the boxes is drawing lines on water.
One version of this result is genuinely humbling. We came in expecting to find that the field had flattened a taxon-structured story into a single number, and we fully intended to be the people who un-flattened it. What the arithmetic says instead is that the taxon-structured story is itself a flattening, imposed on a literature whose disagreements do not respect phylum boundaries at all.
So where does the variance live? Not in taxonomy, or at least not detectably at this sample size. The candidates that remain are not mysterious: experiment duration, whether animals were acclimated or shocked, food supply, the method used to manipulate carbonate chemistry, life stage, and whether the organism lives somewhere that already swings through a wider pH range every day than the experiment imposed on it. Each of those has a literature. Testing them properly would take a few hundred rows and two independent screeners, which is the subject of a later section and not of this one.
The Funnel, and What It Admits
Meta-analysis carries a standard diagnostic for the worry that the published record and the conducted record are not the same record, and it is the funnel plot. Plot each effect against its precision. If everything that was run got published, the cloud should be a symmetric funnel, with precise studies clustered near the pooled value at the top and imprecise ones scattering evenly to both sides below. A gap in one lower corner suggests that small studies finding results in that direction were written up, looked unpublishable to somebody, and never made it into print.
Egger's regression test formalises the worry by regressing each effect on its standard error and asking whether the intercept differs from zero [21], and ours does. The intercept is +0.403 with a standard error of 0.114, giving t = 3.52 on 21 degrees of freedom and p = 0.002. The slope runs negative. Imprecise studies in this collection report systematically more negative effects than the precise ones do, which is the pattern a publication filter would leave behind.
We want to be careful about what that licenses you to conclude. Three things.
Start with power. Egger's test has very little of it below about twenty studies, and it is not reliable when the underlying effects are genuinely heterogeneous, which ours emphatically are. A significant intercept in the presence of I² = 98.8% can be produced by real between-study differences that happen to correlate with study size, with nothing whatever suppressed and no editor having declined anything.
Next, how much of the asymmetry hangs on a single point. Dropping the single most influential row, Iglesias-Rodriguez, flips the intercept from +0.403 to −0.177 while leaving it statistically distinguishable from zero (p = 0.004), so the test stays significant and changes its mind about which way the record is skewed. An asymmetry whose direction reverses when you remove one study is not a finding that anybody should build anything at all on.
The likeliest explanation is duller than suppression. Precision and effect are correlated in this collection for a reason that has nothing to do with what did or did not get published. The most precise rows are cell-culture experiments on single-celled algae, three replicate flasks per treatment, instruments reading to four significant figures; the least precise are whole-animal experiments on urchins and corals, where individual variation is enormous. Those two kinds of experiment measure different biology, and the biology they measure happens to differ in direction. Confounding, then, rather than censorship.
The Strongest Case Against This Article
Now the arguments we would make if we were trying to take this piece apart, made as well as we can make them, because an objection you have not stated properly is an objection you have not answered.
Objection one, and it is the serious one
Your central result is a failure to detect, and you are presenting it as a discovery. QM = 3.14 with p = 0.87 does not mean taxa respond alike; it means 23 effect sizes spread thinly over 8 groups cannot resolve a taxon effect when the residual variance stands at 0.279. Your sample size is talking, not the biology. Run the power calculation for this design and you will find that it could barely detect a taxon difference of the size you yourselves estimated, which makes your headline an absence of evidence dressed as evidence of absence, inside an article whose entire complaint concerns people overreading statistics.
We think this objection is correct, and we have made it easy to check. The interactive companion to this article simulates exactly this design: eight groups, three experiments each, residual variance 0.28, and a true between-taxon spread of 0.48, which is the spread of our own group means. Power comes out at about 24%. Three times in four, a taxon difference of exactly the size we appear to see would go unnoticed by the very test we are relying on. Reaching 80% power on those settings would take roughly 22 experiments per taxon, meaning about 175 in total, against the 23 we have.
So the honest statement is narrower than "taxon does not matter". The honest statement runs: this literature, as we were able to extract it, cannot tell the difference between a world where taxa respond differently and a world where they do not. We will defend that weaker claim rather than the stronger one, and we still think it worth publishing, because the sentence being repeated in public is not "we cannot tell yet". The sentence being repeated is a number.
The reply has a second half. Low power explains why we failed to detect a taxon effect, but it does not explain why QM came out at 3.14 rather than, say, 11, which would still have been non-significant while at least pointing in the direction of a real difference. A statistic of 3.14 against a null expectation of 7 is lower than chance would typically deliver, which is to say that the group means sit closer together, relative to their own uncertainty, than random assignment would produce.
Objection two
Your pooled estimate is meaningless and you should not have computed it. Standard methodological advice, with I² sitting at 98.8%, is that pooling is inappropriate, and you computed the number anyway, put it at the head of your abstract, then spent ten sections instructing readers to ignore it. Having it both ways.
Fair. Our defence is that the pooled number is the thing readers arrive expecting, and that refusing to compute it in an article of this kind would look a great deal like concealment. So we compute it, print its prediction interval directly underneath, and make the gap between the two the entire point of Figure 1. A reader who takes only the abstract away gets −13.0% with an interval containing zero, which is at least honest about how little it knows.
Objection three
Your taxon groups are an artefact of which studies you happened to find. Four of your eight groups rest on one or two experiments. The crustacean row leans on Ries 2009, whose blue crabs and shrimp were grown at 25 degrees with unlimited food, conditions under which almost anything calcifies well. Change the study set and the taxon means change with it, which is presumably the whole reason those means refuse to separate from one another at all.
Also fair, and the leave-one-out analysis is where we test it. Dropping any single row moves the pooled estimate only between −16.6% and −10.3%, so the overall figure is stable, while the taxon means are much less stable and we have not claimed otherwise anywhere above. Notice, though, that this objection strengthens our conclusion rather than weakening it: if the taxon means move around when you swap studies in and out, then the taxon means are not measuring a property of the taxon.
Objection four
Short experiments on adults miss the actual mechanism of harm. Almost everything you pooled is a weeks-long exposure of adult or juvenile organisms. The damage that matters for populations happens at fertilisation, at larval settlement and across generations, and it works through energy budgets rather than through direct dissolution of anything. A crab calcifying faster under acidification may simply be spending energy it needed elsewhere. Your +14.3% could be a cost, misread as a benefit.
This one we cannot answer, and we think it is probably right. Wood, Spicer and Widdicombe demonstrated exactly that pattern in a brittlestar, where calcification and metabolism both rose under acidification and the animal paid for the rise in muscle wastage [17]. Nothing in a log response ratio on calcification rate can detect a trade-off of that kind, which means our numbers describe one measured quantity under one kind of experiment on one life stage, and say nothing whatever about whether crustaceans will be fine.
Objection five
None of this is new and the specialists have known it for years. Kroeker and colleagues pooled 228 studies in 2013 and reported that the differences among taxonomic groups were not a significant source of variation in calcification, which is your conclusion, reached on a far larger evidence base and a decade earlier. You have spent a term rediscovering a subordinate clause.
Correct, and the subordinate clause deserves to be on the record rather than buried.
That 2013 analysis rests on more than fourteen times our evidence base [22]. Its headline figures by taxon look much like ours, with corals, coccolithophores and molluscs showing the largest calcification reductions and no detectable effect for echinoderms or crustaceans. Then comes the sentence we read four times before believing it, reporting that these very differences were not a significant source of variation in the calcification analysis that produced them.
So the result we spent a term arriving at was already in print, stated plainly, in the most-cited meta-analysis in the field, and we did not find it because nobody repeats it. What gets repeated from that paper is the taxon-by-taxon table, which is memorable and visual and fits on a slide, while the sentence saying that the table's rows are not statistically distinguishable sits in a subordinate clause in the results section.
Our contribution, if there is one, is the calibration check in section nine. Kroeker and colleagues reached a non-significant subgroup result and reported it. We reached a wildly significant one using a test that should not have been used, caught it by running that test on data with a known answer, and ended up agreeing with them from the opposite direction. The route matters, because it shows how a real analyst using a standard, widely published procedure can be handed a p-value of 10−158 for an effect that is not there.
What We Are Not, and the Finding Anyway
The limits first, stated flatly, because a club that hides its limits has no business publishing a pooled estimate, and because the limits here are large enough to change how much of the rest you should believe.
We are eleven people with library access and a text editor. A systematic review has two people screening every candidate study independently, a pre-registered protocol, a documented search string run across several databases, and a formal risk-of-bias assessment for each study it includes. We had one search, run by whoever was free that week, conducted mostly through a data repository rather than a literature database, and no independent duplicate screening at any stage of it. No amount of correct arithmetic downstream repairs that, and it is the single largest weakness in the article, larger by some distance than anything in the arithmetic.
What follows from that, concretely. Our 16 studies are not a sample of the acidification literature; they are a sample of the acidification literature that archived usable data in a place we could find it, which biases toward European and North American groups, toward the period after about 2007, and toward organisms that survive in aquaria. Three of our 23 rows have variances we invented. Our grouping decisions, our pCO2 window and our choice of control level were all made by us, in a room, over a few weeks, and several of them could reasonably have gone the other way. The oyster disagreement in section five is a live example: our number and the original authors' number differ by fourteen percentage points because of a binning choice.
So how much weight does −13.0% deserve? Very little as an estimate of anything in the ocean. Rather more as a demonstration that the literature does not support a confident single figure, because a biased sample that still fails to agree with itself is telling you something real about disagreement, whatever it fails to tell you about the sea.
And now the finding.
The variation is not noise to be averaged away, and not an embarrassment to be apologised for in a limitations paragraph; the variation is the signal, and it does not run where everybody assumes it runs. Ninety-nine percent of the differences between these experiments are real rather than sampling error, and almost none of that is explained by which organism happened to be sitting in the tank when the pH came down. Two coral experiments differ about as much as a coral and a crab.
A more uncomfortable result than "some taxa are hit and others are not", and a more useful one, because if taxonomy explained the spread the research programme would simply be to work through the phyla one at a time. Since it does not, the programme has to be about the experiments themselves: how long they ran, what the animals were fed, how the carbonate chemistry was manipulated, what life stage went into the tank, and how the organism's home habitat compares with the treatment it was given. Those axes are testable. No pooled percentage shows any of them, and neither does a taxon-by-taxon table, which is why we would rather leave you with a question than a figure.
Here is the plate one last time, drawn the way the arithmetic actually leaves it, with every group open, because no group is separated from the others by more than the noise inside it, and with each cell carrying its own within-group I² in place of the verdict we spent most of a term wanting to give it.
k = 7
k = 2
k = 3
k = 2
k = 3
k = 2
k = 3
The title of this piece is a small provocation and we will still defend it, though not quite in the sense we meant when we chose it. Not every shell dissolves. Some do, some thin, some thicken, and one urchin got heavier. Saying so is not scepticism about acidification; it is the ordinary business of reading a literature carefully enough to notice that it contains more than one result, and then carefully enough again to notice that the results do not line up along the axis the story requires them to line up along.
References
- Riebesell, U., Zondervan, I., Rost, B., Tortell, P. D., Zeebe, R. E. & Morel, F. M. M. (2000). Reduced calcification of marine plankton in response to increased atmospheric CO2. Nature 407, 364–367. doi:10.1038/35030078. Data: PANGAEA doi:10.1594/PANGAEA.728092
- Gazeau, F., Quiblier, C., Jansen, J. M., Gattuso, J.-P., Middelburg, J. J. & Heip, C. H. R. (2007). Impact of elevated CO2 on shellfish calcification. Geophysical Research Letters 34, L07603. doi:10.1029/2006GL028554
- Iglesias-Rodriguez, M. D., Halloran, P. R., Rickaby, R. E. M., Hall, I. R., Colmenero-Hidalgo, E., Gittins, J. R., Green, D. R. H., Tyrrell, T., Gibbs, S. J., von Dassow, P., Rehm, E., Armbrust, E. V. & Boessenkool, K. P. (2008). Phytoplankton calcification in a high-CO2 world. Science 320, 336–340. doi:10.1126/science.1154122. Data: PANGAEA doi:10.1594/PANGAEA.718841
- Comeau, S., Gorsky, G., Jeffree, R., Teyssié, J.-L. & Gattuso, J.-P. (2009). Impact of ocean acidification on a key Arctic pelagic mollusc (Limacina helicina). Biogeosciences 6, 1877–1882. doi:10.5194/bg-6-1877-2009
- Maier, C., Hegeman, J., Weinbauer, M. G. & Gattuso, J.-P. (2009). Calcification of the cold-water coral Lophelia pertusa under ambient and reduced pH. Biogeosciences 6, 1671–1680. doi:10.5194/bg-6-1671-2009
- Langer, G., Nehrke, G., Probert, I., Ly, J. & Ziveri, P. (2009). Strain-specific responses of Emiliania huxleyi to changing seawater carbonate chemistry. Biogeosciences 6, 2637–2646. doi:10.5194/bg-6-2637-2009
- Ries, J. B., Cohen, A. L. & McCorkle, D. C. (2009). Marine calcifiers exhibit mixed responses to CO2-induced ocean acidification. Geology 37, 1131–1134. doi:10.1130/G30210A.1. Data: PANGAEA doi:10.1594/PANGAEA.733947
- Findlay, H. S., Kendall, M. A., Spicer, J. I. & Widdicombe, S. (2010). Relative influences of ocean acidification and temperature on intertidal barnacle post-larvae at the northern edge of their geographic distribution. Estuarine, Coastal and Shelf Science 86, 675–682. doi:10.1016/j.ecss.2009.11.036
- Lischka, S., Büdenbender, J., Boxhammer, T. & Riebesell, U. (2011). Impact of ocean acidification and elevated temperatures on early juveniles of the polar shelled pteropod Limacina helicina: mortality, shell degradation, and shell growth. Biogeosciences 8, 919–932. doi:10.5194/bg-8-919-2011
- Chauvin, A., Denis, V. & Cuet, P. (2011). Is the response of coral calcification to seawater acidification related to nutrient loading? Coral Reefs 30, 911–923. doi:10.1007/s00338-011-0786-7
- Uthicke, S. & Fabricius, K. E. (2012). Productivity gains do not compensate for reduced calcification under near-future ocean acidification in the photosynthetic benthic foraminifer species Marginopora vertebralis. Global Change Biology 18, 2781–2791. doi:10.1111/j.1365-2486.2012.02715.x
- Courtney, T., Westfield, I. & Ries, J. B. (2013). CO2-induced ocean acidification impairs calcification in the tropical urchin Echinometra viridis. Journal of Experimental Marine Biology and Ecology 440, 169–175. doi:10.1016/j.jembe.2012.11.013
- Long, W. C., Swiney, K. M., Harris, C., Page, H. N. & Foy, R. J. (2013). Effects of ocean acidification on juvenile red king crab (Paralithodes camtschaticus) and Tanner crab (Chionoecetes bairdi) growth, condition, calcification, and survival. PLoS ONE 8, e60959. doi:10.1371/journal.pone.0060959
- Schoepf, V., Grottoli, A. G., Warner, M. E., Cai, W.-J., Melman, T. F., Hoadley, K. D., Pettay, D. T., Hu, X., Li, Q., Xu, H., Wang, Y., Matsui, Y. & Baumann, J. H. (2013). Coral energy reserves and calcification in a high-CO2 world at two temperatures. PLoS ONE 8, e75049. doi:10.1371/journal.pone.0075049
- Comeau, S., Cornwall, C. E., DeCarlo, T. M., Krieger, E. & McCulloch, M. T. (2018). Similar controls on calcification under ocean acidification across unrelated coral reef taxa. Global Change Biology 24, 4857–4868. doi:10.1111/gcb.14379. Data: PANGAEA doi:10.1594/PANGAEA.892655
- Edmunds, P. J. & Burgess, S. C. (2020). Emergent properties of branching morphologies modulate the sensitivity of coral calcification to high PCO2. Journal of Experimental Biology 223, jeb217000. doi:10.1242/jeb.217000. Data: PANGAEA doi:10.1594/PANGAEA.926042
- Wood, H. L., Spicer, J. I. & Widdicombe, S. (2008). Ocean acidification may increase calcification rates, but at a cost. Proceedings of the Royal Society B 275, 1767–1773. doi:10.1098/rspb.2008.0343
- Hedges, L. V., Gurevitch, J. & Curtis, P. S. (1999). The meta-analysis of response ratios in experimental ecology. Ecology 80, 1150–1156. doi:10.1890/0012-9658(1999)080[1150:TMAORR]2.0.CO;2
- DerSimonian, R. & Laird, N. (1986). Meta-analysis in clinical trials. Controlled Clinical Trials 7, 177–188. doi:10.1016/0197-2456(86)90046-2
- Higgins, J. P. T. & Thompson, S. G. (2002). Quantifying heterogeneity in a meta-analysis. Statistics in Medicine 21, 1539–1558. doi:10.1002/sim.1186
- Egger, M., Davey Smith, G., Schneider, M. & Minder, C. (1997). Bias in meta-analysis detected by a simple, graphical test. BMJ 315, 629–634. doi:10.1136/bmj.315.7109.629
- Kroeker, K. J., Kordas, R. L., Crim, R., Hendriks, I. E., Ramajo, L., Singh, G. S., Duarte, C. M. & Gattuso, J.-P. (2013). Impacts of ocean acidification on marine organisms: quantifying sensitivities and interaction with warming. Global Change Biology 19, 1884–1896. doi:10.1111/gcb.12179