Science Journaling Club Founded 2024

REVIEW · META-ANALYSIS · SOIL SCIENCE

How Much Carbon Can Farmland Actually Hold? A Review of the Field Trial Evidence

Written jointly by the Science Journaling Club

Review · Peer-edited by the club review board · LaTeX source · Our calculation · Interactive model

Abstract Farmers are being offered money to store carbon in their soil, the practices on the menu being no-till, cover cropping and residue retention, and the rates written into contracts usually sit somewhere near half a tonne of carbon per hectare per year. We went looking for where that number comes from. Screening 22 published syntheses and long-term trial reports, we could extract 12 effect sizes carrying usable uncertainty from six publications, then pooled them with a random-effects model we wrote ourselves and checked against synthetic data first. The pooled increase in soil organic carbon is +16.2% (95% CI 11.3 to 21.2), and nobody should use that number for anything, because \(I^2 = 98\%\) and the 95% prediction interval for a new field runs from −1.7% to +37.3%. Two results survive the heterogeneity. Practices that import carbon onto the farm hold their effect below the plough layer while practices that only rearrange carbon already there lose it, and pooled, the two groups differ by a factor of 3.3 (\(p < 10^{-29}\)). In a single dataset of 95 paired comparisons, switching from fixed-depth accounting to equal-soil-mass accounting cuts the measured no-till sequestration rate from 0.300 to 0.141 Mg C ha⁻¹ yr⁻¹, a factor of 2.1, with nothing whatever happening in the field [1]. Soils also saturate, and a one-line model fitted to no data and needing none says that a practice adding 0.5 Mg C ha⁻¹ yr⁻¹ of input buys about 53 years before the annual gain drops below a fifth of what it started at, and then stops. We are a school club reading abstracts rather than a systematic review with two independent screeners, and §3 says exactly what that costs.

Section 1 · surface, 0–5 cm

The Number That Made Us Suspicious

Half a tonne per hectare per year.

Roughly the figure a farmer sees when a soil carbon programme comes calling, roughly the figure we came to check, and on its face nothing about it looks outrageous. Multiply it by the 44/12 conversion from carbon to carbon dioxide and you get about 1.8 tonnes of CO₂ equivalent per hectare per year, which at fifty dollars a tonne is ninety dollars a hectare, which on a decent-sized arable holding is real money arriving for doing something you were half-minded to do anyway. The pitch writes itself: stop ploughing, plant a cover crop over winter, leave the straw where it falls, drill straight into the residue next spring, and the soil quietly banks carbon that somebody else pays you for.

What made us suspicious was not the agronomy, which is sound and in places genuinely elegant. Our suspicion came from watching the same number surface again and again in documents whose authors had every commercial reason to want it large, and surface without the two things any measured quantity should carry with it: how deep they dug, and how long they waited.

So we went and looked, and found that the whole estimate turns on the depth question, which nobody had told us was anything more than a technicality. No-till does something real and well documented to a soil profile, concentrating organic matter near the surface, so that a survey sampling only near the surface measures a gain the whole profile does not have. Soil science has known this since at least 2007 [19]. The market has been slower to catch up.

This article is a meta-analysis, so the sources are the data: everything we did sits in the script, every number we extracted sits in §4 with a note saying which sentence of which paper it came out of, and §3 gives an honest account of why our pooled figure deserves less weight than the individual long-term experiments underneath it. Those experiments are the real heroes here, several of them running since the nineteenth century on funding nobody enjoys renewing, and they are the only reason any of this is arguable rather than merely assertable.

Section 2 · plough layer, 0–30 cm

What a Soil Carbon Credit Is, Mechanically

A credit is a promise about a quantity, so start with the quantity. Soil organic carbon, SOC, is the carbon held in dead and decaying plant and microbial material in the soil, as distinct from the carbonate minerals that also contain carbon but do not respond to farming on any timescale that matters here. Laboratories report it one of two ways: as a concentration, grams of carbon per kilogram of soil, which is what a combustion analyser actually measures, or as a stock, tonnes of carbon per hectare, which is the only one of the two that a market can sell.

Turning the first into the second requires the bulk density, the mass of a given volume of soil, and there the trouble starts, because a stock is concentration times bulk density times depth. Sample a 30 centimetre core, measure 12 g C per kg and a bulk density of 1.30 g/cm³, and you have 12 × 1.30 × 30 × 10 = 4,680 g C per square metre, or 46.8 tonnes per hectare. Clean arithmetic.

Now stop ploughing that field for fifteen years. Earthworms return and aggregates stabilise, the soil loosens, and the bulk density in the top layer falls to 1.20, so your 30 centimetre core now holds less soil than it used to. Hold the carbon concentration exactly constant and the calculated stock has still fallen; let the concentration rise a little and you may record a gain, a loss or nothing at all, depending entirely on which of the two changes won.

Soil scientists fixed this decades ago by comparing equivalent soil masses rather than equivalent depths: dig deeper in the loosened soil until the sample holds the same mass of mineral matter, then compare. Raffeld and colleagues, working on twenty years of data from two long-running American cropping trials, found that fixed-depth and equivalent-soil-mass estimates of stock change can differ by over a hundred per cent, and that the difference tracks the bulk-density change in the surface soil almost perfectly, with a correlation of 0.90 in one set of treatments [22]. They recommend equivalent soil mass, sampled to 60 cm in multiple increments. Most market protocols require neither.

Section 3 · no depth, this is about us

What We Did, and Everything That Is Wrong With It

We searched for published syntheses and multi-site analyses reporting soil organic carbon change under no-till, reduced tillage, cover cropping, crop residue or straw retention and crop rotation, and we deliberately also collected estimates for manure and for organic amendments substituted for mineral fertiliser, because those practices physically import carbon onto the field and so act as a control group for the central question. If a practice that genuinely adds carbon behaves one way with depth while a practice that only moves carbon around behaves another, the difference between them is evidence about mechanism rather than only about magnitude.

One person extracted each number and a second checked it against the source text, and every extracted value is quoted in §4 alongside the sentence it came out of.

Now the list of what is wrong with that, in descending order of how much it should worry you.

Our entries are not independent. This is the serious one. Published syntheses of the same literature share primary studies, sometimes heavily, and every pooling formula in the script assumes the effect sizes are independent draws, which ours plainly are not. Beillouin and colleagues explicitly report the fraction of primary studies used by more than one meta-analysis in their own dataset, which tells you they know the problem is real [2]. We report a pooled number because the brief asked for one and because the spread it exposes is informative. Nobody should price anything off it.

We screened abstracts rather than papers. A systematic review has a registered protocol, two people screening every record independently, a documented disagreement-resolution procedure and attempts to reach authors for unpublished data, and we have none of that. What we had was a search engine, an institutional-free open-access route to full texts when the publisher allowed it, and a Tuesday afternoon.

The discard pile is bigger than the pool, and it is not a random sample. We extracted 27 effect sizes and could pool 12, and what got discarded skews toward nulls, because a null is often reported as the bare sentence "no significant effect was detected" with no point estimate at all, and toward absolute stock differences in tonnes per hectare, which cannot be converted to a proportional change without a control stock the paper never gave. Both of those categories are systematically smaller than what survived, so our pooled estimate is biased upward by our own inclusion rule, and we cannot tell you by how much.

Two of our twelve intervals are derived rather than reported. Liu and colleagues give "12.8 ± 0.4%" without saying whether the ± is a standard error or a confidence interval, so we read it as a standard error and widened it [4]; Han and colleagues give an absolute change of 2.0 g/kg with an interval of 1.9 to 2.2 alongside a relative change of 19.5%, so we scaled the absolute interval by the ratio [5]. Both are flagged in the table, and both are dropped in a sensitivity run in §10.

A club extracting effect sizes from abstracts is not a systematic review, and we would rather say so in the third section than bury it in the twelfth.

Section 4 · the whole sampled column

The Ledger

Here is every number, where it came from and whether we could use it, the pool in the top block and the discard pile in the bottom one, the latter listed in full because a review that shows you only its survivors is showing you a selected sample and calling it a census.

Key Source Practice Effect 95% interval Reported n Max depth Provenance
Pooled · proportional change in soil organic carbon
Du17-ESMDu et al. [1]no-till, equivalent soil mass+3.8%1.4 to 6.395 comparisons, 57 sites40 cmabstract, verbatim
Du17-FDDu et al. [1]no-till, fixed depth+5.1%2.5 to 7.795 comparisons, 57 sites30 cmabstract, verbatim
Bei23-NTBeillouin et al. [2]no-till+9.3%5.6 to 13.0230 meta-analysesnot statedresults text, verbatim
Bei23-RTBeillouin et al. [2]reduced tillage+12.0%0.1 to 24.0second-ordernot statedresults text, verbatim
Bei23-RESBeillouin et al. [2]crop residue retention+13.0%9.8 to 16.0second-ordernot statedresults text, verbatim
Bei23-ROTBeillouin et al. [2]crop rotation+6.5%−0.9 to 14.0second-ordernot statedresults text, verbatim
Jian20-CCJian et al. [3]cover cropping+15.5%13.8 to 17.3not stated in abstract30 cmabstract, verbatim
Liu14-STRLiu et al. [4]straw return+12.8%12.0 to 13.6176 field studies20 cmderived from ±0.4
Han16-CFSHan et al. [5]straw plus fertiliser+19.5%18.5 to 21.5global compilation20 cmscaled from g/kg
Han16-CFMHan et al. [5]manure plus fertiliser+36.2%33.1 to 39.3global compilation20 cmscaled from g/kg
GG21-MANGross & Glaser [6]manure+35.0%32.0 to 39.0592 comparisons, 101 studies30 cmresults text, verbatim
Bei23-ORGBeillouin et al. [2]organic for mineral fertiliser+34.0%20.0 to 49.0second-ordernot statedresults text, verbatim
Pooled separately · annual accrual rate, Mg C ha⁻¹ yr⁻¹
WP02West & Post [17]no-till0.57±0.14 (s.e.)276 paired treatmentsmostly 30 cm57 ± 14 g C m⁻² yr⁻¹
PD15Poeplau & Don [18]cover crops0.32±0.08 (s.e.)139 plots, 37 sites22 cm meanabstract, verbatim
Du17-FDrDu et al. [1]no-till, fixed depth0.3000.054 to 0.54795 comparisons30 cmabstract, verbatim
Du17-ESMrDu et al. [1]no-till, equivalent soil mass0.141−0.102 to 0.38495 comparisons30 cmabstract, verbatim
Extracted, read, and not poolable
 Luo et al. [7]no-till, 0–10 cm+3.15 Mg/ha±2.4269 paired experiments10 cmstock, no control stock
 Luo et al. [7]no-till, 20–40 cm−3.30 Mg/ha±1.6169 paired experiments40 cmstock, no control stock
 Luo et al. [7]no-till, 0–40 cmno increasenone given69 paired experiments40 cmreported as a null
 Haddaway et al. [8]no-till, 0–30 cm, ≥10 yr+4.6 Mg/ha0.78 to 8.43boreo-temperate30 cmstock, no control stock
 Haddaway et al. [8]no-till, full profileno effectnone givenboreo-temperatefullreported as a null
 Angers & Eriksen-Hamel [9]no-till vs inversion tillage+4.9 Mg/hanot retrievedpaired, ≥5 yr, ≥30 cmprofileno interval available
 Virto et al. [10]no-till, 0–30 cm+6.7%not reported92 paired cases30 cmno interval in abstract
 McClelland et al. [11]cover crops, 0–30 cm+12%figure only181 obs, 40 papers30 cminterval in a figure
 Prairie et al. [12]no-till, 0–20 cm+11.3%figure only118 studies20 cminterval in Fig. 2
 Prairie et al. [12]no-till, below 20 cmnot significantnone given118 studies>20 cmreported as a null
 Bai et al. [13]conservation tillage+5%not reported3,049 paired measurementsmixedno interval in abstract
 Bai et al. [13]cover crops+6%not reported3,049 paired measurementsmixedno interval in abstract
 Crystal-Ornelas et al. [14]conservation tillage, organic+14%not reportedorganic systemsweightedno interval in abstract
 Emde et al. [15]irrigation+5.9%figure only2–47 yr studiesmixedinterval in Fig. 3
 Sun et al. [16]no-tillregional onlynone given260 paired studiesmixedno global summary

Twenty-two sources read. Twenty-seven effect sizes extracted. Twelve poolable, from six publications. Those three numbers are the honest description of this review, and any reader who stops here has the important part.

Section 5 · 0–30 cm, where the studies are

Pooling, and Why the Pooled Number Should Not Be Used

We wrote the estimator from the formulas rather than calling a package, partly so that the club could see what the arithmetic actually does and partly because a routine you have not checked is a routine you have no business trusting. Effects are log response ratios, \(y = \ln(1 + p/100)\) for a percentage change \(p\), because response ratios multiply rather than add and their sampling distribution is far closer to normal in logs. Between-study variance is estimated by the DerSimonian-Laird moment estimator,

$$\hat\tau^2 = \max\left(0,\ \frac{Q - (k-1)}{\sum w_i - \sum w_i^2 / \sum w_i}\right), \qquad Q = \sum_i w_i (y_i - \hat\theta_{\text{FE}})^2$$

and the random-effects weights are \(w_i^* = 1/(v_i + \hat\tau^2)\). Before running it on anything real we made it recover an answer we already knew: twelve synthetic studies drawn from a true effect of exactly +12% with a true \(\tau^2\) of 0.0025 came back at +12.09% with a 95% interval of 8.72 to 15.56, which contains the truth. Repeating the whole exercise 2,000 times, the mean estimated \(\tau^2\) was 0.002498 against a true 0.0025, while interval coverage came out at 91.5% rather than the nominal 95%, a shortfall that is the known behaviour of this estimator rather than a fault in our code, since it treats \(\hat\tau^2\) as though it were known exactly when it is nothing of the kind. Our intervals on the real data are therefore slightly too narrow, which we would rather print than hide.

Second check, and this one is an algebraic identity rather than a simulation. Hand the routine a dataset with no between-study heterogeneity at all and \(\hat\tau^2\) must come out at zero; at \(\hat\tau^2 = 0\) the random-effects weight \(1/(v_i + \tau^2)\) is the fixed-effect weight \(1/v_i\), so the two weighted means must be the same weighted mean. They are, to a difference of exactly 0.000e+00.

Now the real data, drawn as Figure 1.

0 10 20 30 40 50 change in soil organic carbon (%) study / practice effect Du17-ESM no-till, ESM+3.8 Du17-FD no-till, fixed depth+5.1 Bei23-NT no-till+9.3 Bei23-RT reduced tillage+12.0 Bei23-RES residue retention+13.0 Bei23-ROT crop rotation+6.5 Jian20-CC cover crops+15.5 Liu14-STR straw return+12.8 Han16-CFS straw + NPK+19.5 Han16-CFM manure + NPK+36.2 GG21-MAN manure+35.0 Bei23-ORG organic substitution+34.0 POOLED (random effects) +16.2 95% prediction interval I² = 98.1% Q = 571.0 on 11 df τ² = 0.0051
Figure 1. Forest plot of the twelve poolable effect sizes, boxes scaled by random-effects weight, whiskers showing the 95% confidence interval each source reported. Dark whiskers are practices that rearrange carbon already in the field; pale whiskers are practices that import it. The diamond is the pooled random-effects estimate, +16.2% (95% CI 11.3 to 21.2). The dashed bar beneath it is the 95% prediction interval, −1.7% to +37.3%, which is the range a new field would plausibly land in and the only summary here we would defend. Numbers from soil-carbon-review.py, Part 3.

Read the bottom bar, not the diamond: \(Q = 571.0\) on 11 degrees of freedom gives \(I^2 = 98.1\%\), which means essentially none of the spread between these studies is sampling noise. These are twelve measurements of twelve different quantities rather than twelve noisy measurements of one, taken in different soils, under different climates, over different durations and to different depths, so calling their weighted average "the effect of conservation agriculture on soil carbon" is a category error dressed up as a statistic.

The weights make the same point from a different angle: with \(\tau^2\) at 0.0051 and sampling variances two orders of magnitude smaller, \(\tau^2\) dominates every denominator and the weights flatten out to between 5.8% and 9.2%, so that a synthesis of 592 paired comparisons counts for about as much as one sentence lifted from a second-order review. Random-effects weighting behaves exactly that way when heterogeneity is enormous, the behaviour is correct, and it should still make you uncomfortable.

Section 6 · arguing, no fixed depth

Working Notes From the Club Table

Session one · the scope argument

R: if we include manure this stops being a review of carbon farming and becomes a review of everything.

K: manure is the control. If no-till and manure both show a topsoil gain and only manure keeps it at 40 cm, we've learned something about mechanism. Without manure we've got a list.

R: fine but label it. Don't let it sit in the same pool unmarked.

Agreed: two families, move and add, flagged in the data structure, pooled together once and separately once.

Session two · the Beillouin problem

Four of our twelve rows come out of one paper [2]. K says that's double-counting and we should take one. R says it's a second-order review so its rows are genuinely different practices and the alternative is a pool of nine.

Compromise nobody loved: keep all four, then run a pre-planned one-row-per-publication sensitivity. It comes back at +15.6% on six rows against +16.2% on twelve. So the double-counting is not what is driving the answer. The heterogeneity is.

Session three · the sentence that changed the article

Somebody read the Du abstract out loud [1]. Same 95 comparisons. Same 57 sites. Fixed depth gives 0.300 Mg C ha⁻¹ yr⁻¹ and it is significant. Equivalent soil mass gives 0.141 and it is not.

K: so the difference between a credit and no credit is which arithmetic you use.

Long pause. Then twenty minutes of somebody insisting it must be a typo. The number is printed in the abstract, twice, and both figures carry their own confidence intervals.

We restructured the article around that sentence, which is now §8.

Session four · the thing we could not fix

Only seven of our twelve rows state a maximum sampling depth at all. Seven. In a literature whose central methodological dispute is about sampling depth.

R: that's not a limitation of our review, that's a finding about the field.

We wrote it up as one, and it opens §7.

Session five · on being annoying about it

Note from the review board, verbatim: "You are three drafts into sounding like people who think farmers are being fooled. They are not being fooled. They are being offered a contract with a measurement clause they did not write and cannot audit. Fix the tone."

Fixed, we think. The scepticism belongs to the accounting, not to anybody standing in a field.

Section 7 · 30–60 cm, below the plough

The Depth Problem

Seven of our twelve pooled effect sizes state a maximum sampling depth, and the other five do not state one anywhere we could find. The first result of this section took us by surprise, and it is worse than it sounds, because a synthesis that does not report its depth range has usually not restricted it, which in practice means it is dominated by studies that sampled the top 20 or 30 centimetres, that being what most trials did.

Here is what happens to a soil profile under no-till: residues stay on the surface instead of being turned under, so organic matter accumulates in the top few centimetres and the profile becomes strongly stratified, while the layer the plough used to enrich, the bottom of the old furrow slice around 20 to 30 centimetres, stops receiving buried residue and loses carbon. Luo and colleagues measured exactly this across 69 paired experiments where sampling went below 40 centimetres: a gain of 3.15 ± 2.42 Mg C ha⁻¹ in the surface 10 centimetres, a loss of 3.30 ± 1.61 Mg C ha⁻¹ in the 20 to 40 centimetre layer, and no increase in total carbon down to 40 centimetres [7]. Figure 2 is that result at scale.

0–10 cm 10–20 cm 20–40 cm +3.15 ± 2.42 Mg C/ha not reported separately −3.30 ± 1.61 Mg C/ha −5 −3 −1 0 +1 +3 +5 change in soil carbon stock, Mg C per hectare NET, 0–40 cm: no increase in total soil carbon. 69 paired experiments · Luo et al. 2010 [7]
Figure 2. What no-till does to a soil profile. The surface gains, the layer below the old furrow slice loses, and the total down to 40 centimetres does not move. A 10 centimetre core would have recorded a handsome gain; a 40 centimetre core recorded nothing. The dashed box is the 10 to 20 cm layer, which the source does not report separately. Values as published [7].

That pattern is not one team's result; it recurs whenever anybody looks. Prairie, King and Cotrufo, synthesising 118 studies and 157 experiments, found no-till raised topsoil carbon by 11.3% from 0 to 20 centimetres and had no significant effect below 20 [12]. Jian and colleagues found cover cropping significantly raised carbon in soils at or above 30 centimetres and not below [3], while Haddaway and colleagues found 4.6 Mg ha⁻¹ in the top 30 centimetres over comparisons of at least ten years and no effect whatever across the full profile [8]. Angers and Eriksen-Hamel, restricting themselves to replicated randomised paired trials of at least five years measured to at least 30 centimetres, found the surface gain under no-till was largely offset by greater carbon near the bottom of the plough layer under inversion tillage [9].

Now the control. Gross and Glaser pooled 592 paired comparisons of manure application, which is carbon physically carted onto the field from somewhere else, and found a 40% increase where sampling stopped at 15 centimetres against a 23% increase below 30 centimetres [6]. Smaller at depth, certainly, but present, significant, and not reversing sign.

Five practices that move carbon lose the effect below the plough layer. One practice that brings carbon in keeps it. The whole argument fits in two sentences.

We tried to test this as a meta-regression of effect size on maximum sampling depth across the seven rows that state one, and the slope comes out at −0.0071 in log units per centimetre, which works out at −6.9% per additional ten centimetres of required depth, with \(p = 0.20\). We are not going to dress that up. With seven points and exactly one beyond 30 centimetres the regression is underpowered, and the honest statement is that the between-study evidence is too thin to test the thing the field most needs tested, which is why the within-study paired contrasts in the paragraphs above carry far more weight here: they hold soil, climate, crop and duration constant and vary only depth.

One calculation makes the stakes visible, and Figure 3 draws it. Suppose a credit protocol refused to count any study that had not sampled to at least a given depth: what survives, and what does the remainder say?

0 5 10 15 20 25 30 pooled SOC change (%) any depth ≥ 20 cm ≥ 30 cm ≥ 40 cm k = 12 k = 7 k = 4 k = 1 +16.2 +17.7 +14.2 +3.8 minimum sampling depth a protocol would require whiskers are 95% confidence intervals
Figure 3. The cost of asking for a deeper core. Each point is the pooled random-effects estimate over the effect sizes that would still qualify under a minimum-depth rule. Requiring 40 centimetres leaves exactly one qualifying entry in our whole dataset, and that entry, which is also the only one using equivalent-soil-mass accounting, reports +3.8% (95% CI 1.4 to 6.3) [1]. The single point is plotted in a different colour because \(k = 1\) is not a pooled estimate, it is a study.

That last point is the article in one mark on a chart. Demand that a study sample to 40 centimetres and correct for bulk density, and our entire twelve-entry dataset reduces to one number, which is four times smaller than the pooled estimate and comes with the observation, from the same authors in the same abstract, that the effect on the annual rate is not statistically distinguishable from zero [1].

Section 8 · 0–30 cm, the core that was taken

Plain Arithmetic

This section states numbers. Du and colleagues analysed 95 paired comparisons from 57 experimental sites in China, all running three years or more [1]. They computed the no-till effect on the top 30 centimetres two ways.

fixed depth
  rate 0.300 Mg C ha⁻¹ yr⁻¹   95% CI 0.054 to 0.547   p < 0.001

equivalent soil mass
  rate 0.141 Mg C ha⁻¹ yr⁻¹   95% CI −0.102 to 0.384   p = 0.100

ratio   0.300 / 0.141 = 2.13
difference   0.300 − 0.141 = 0.159 Mg C ha⁻¹ yr⁻¹

Same fields. Same cores. Same laboratory. The second figure corrects for no-till soil being less dense, so a fixed 30 centimetre core holds less soil under no-till than under the plough.

Carbon to carbon dioxide is 44/12 = 3.667.

0.300 × 3.667 = 1.100 t CO₂e ha⁻¹ yr⁻¹
0.141 × 3.667 = 0.517 t CO₂e ha⁻¹ yr⁻¹

at $50 per t CO₂e
  fixed depth         $55.00 per hectare per year
  equivalent soil mass   $25.85 per hectare per year

200 hectares × $25.85 = $5,170 per year, gross of sampling, laboratory, verification and registry fees

For comparison, the two rate estimates most often cited. West and Post, from 67 long-term experiments and 276 paired treatments, give 57 ± 14 g C m⁻² yr⁻¹ for a change from conventional tillage to no-till, which is 0.57 ± 0.14 Mg C ha⁻¹ yr⁻¹ [17]. Poeplau and Don, from 139 plots at 37 sites, give 0.32 ± 0.08 Mg C ha⁻¹ yr⁻¹ for cover crops at a mean sampling depth of 22 centimetres [18].

four rate estimates pooled, random effects
  0.324 Mg C ha⁻¹ yr⁻¹   95% CI 0.175 to 0.473
  Q = 5.30 on 3 df, τ² = 0.0100, I² = 43.4%

= 1.186 t CO₂e ha⁻¹ yr⁻¹ = $59.32 per hectare per year at $50

The heterogeneity in this small pool is 43.4%, which is moderate rather than catastrophic, and the reason is that all four estimates come from the same shallow depth band. Restrict the question enough and the studies start agreeing. Restricting the question is not the same as answering it.

Section 9 · whole profile, over decades

The Ceiling

Even if every number above were larger, there would be a second problem, and it is the one that long-term experiments are uniquely placed to see. Soil carbon does not accumulate forever.

Write the simplest model anybody could write: carbon enters the stable soil pool at some annual rate \(I\) set by plant residues and roots, and leaves through microbial respiration at a rate proportional to how much is already there, with fractional loss constant \(k\). Then

$$\frac{dC}{dt} = I - kC, \qquad C(t) = C_{\text{eq}} + (C_0 - C_{\text{eq}})e^{-kt}, \qquad C_{\text{eq}} = \frac{I}{k}$$

Raising the input raises the equilibrium in exact proportion, and the approach to that equilibrium is exponential. The practice does not stop working because the farmer stops trying; it stops working because the stock has climbed close enough to its new ceiling that respiration has caught up with the extra input, and by then there is nothing left to sell. Figure 4 draws one soil doing it.

ceiling C_eq = 61.7 Mg C/ha year 53.6: annual gain falls below 0.10 40 45 50 55 60 65 stock, Mg C per hectare 0 0.25 0.50 0 50 100 150 years since the practice started stock (left axis) annual gain (right axis) C₀ = 45, k = 0.03/yr, extra input 0.50 Mg C/ha/yr
Figure 4. A soil filling up. With a starting stock of 45 Mg C ha⁻¹ in the top 30 centimetres, a loss constant of 0.03 per year and a practice adding 0.50 Mg C ha⁻¹ yr⁻¹ of input, the total gain available is 16.7 Mg C ha⁻¹ and not one tonne more. Half of it arrives by year 23, ninety per cent by year 77, and the annual gain has fallen below a fifth of its starting value by year 53.6. Change the settings yourself in the interactive model.

Nothing in that model was fitted to anything: it carries two constants and one line of calculus, and the published durations line up with it uncomfortably well. Liu and colleagues, from 176 field studies, put saturation under straw return at around twelve years [4], and Han and colleagues estimate sequestration durations of 28 to 73 years for straw plus fertiliser and 26 to 117 years for manure plus fertiliser, with wide variation across climates [5]. Poeplau and Don, running their cover-crop data through a turnover model, get a new steady state after roughly 155 years with a total accumulation of 16.7 ± 1.5 Mg C ha⁻¹ at 22 centimetres depth, within a rounding error of the ceiling our two-constant model produces by accident [18].

The strongest evidence on ceilings comes from the experiments nobody wants to fund. Poulton and colleagues went through 114 treatment comparisons across 16 long-term experiments at Rothamsted, some running for 157 years, and asked whether the "4 per 1000" target of a 4 per mille annual increase in soil carbon was achievable [20]. In the two longest experiments, annual farmyard manure at 35 tonnes fresh weight per hectare produced increases of 18 and 43 per mille per year over the first twenty years, and rates above 7 per mille continued for another forty to sixty. Then they fell away. A soil receiving an enormous and unbroken carbon input from outside still slows down, because it is filling.

Georgiou and colleagues put a mechanism under this from 1,144 soil profiles worldwide: carbon bound to mineral surfaces is the durable fraction, mineral surfaces are finite, and soils currently sit at about 42% of their mineralogical capacity in surface layers and 21% at depth [21]. Across 103 carbon-accrual measurements they found that soils furthest from capacity accrue fastest, with rates averaging three times higher at a tenth of capacity than at a half, while our one-line model, given a 100 Mg ha⁻¹ capacity, produces a ratio of 4.4 for the same comparison. Right sign, roughly right size, out of a model told nothing whatever about minerals, which we take as encouragement rather than as confirmation.

Section 10 · turning the instrument on ourselves

Our Own Funnel, and What It Cannot Tell Us

A funnel plot puts each study's effect against its precision, so that precise studies cluster near the pooled value at the top while imprecise ones scatter symmetrically below, and the shape should be a funnel. When the bottom left of that funnel comes out empty, the usual reading is that small studies finding small or negative effects were run and never published. Figure 5 is ours.

Liu14 Han16-CFS Jian20 Han16-CFM GG21 Du17 Bei23-NT Bei23-ROT Bei23-RT Bei23-ORG 0 0.10 0.20 0.30 effect size, log response ratio 0 0.02 0.04 0.06 standard error Egger intercept +1.61 (s.e. 3.56), t = 0.45, p = 0.66
Figure 5. Funnel plot of the twelve effect sizes. The shaded wedge is the pseudo 95% funnel around the pooled estimate. Almost nothing sits inside it, which is what \(I^2 = 98\%\) looks like drawn as a picture. The two clusters are the two families: rearranging practices on the left, importing practices on the right. Egger's test returns an intercept of +1.61 with a standard error of 3.56, which is a way of saying the test could not see anything through the heterogeneity.

Our funnel does not fail that test so much as decline to take it: the points form two clusters at different effect sizes rather than a funnel, because the dataset contains two mechanistically different kinds of practice. Egger's regression of the standard normal deviate on precision gives an intercept of +1.61 against a standard error of 3.56, so \(t = 0.45\) on 10 degrees of freedom and \(p = 0.66\). With twelve points and \(I^2\) at 98% that test has almost no power, and heterogeneity manufactures funnel asymmetry on its own even where no publication bias exists, so we report the number because we said we would, and nobody should read it as evidence of absence.

Leave-one-out is more informative: removing each study in turn moves the pooled estimate over a range of 14.3% to 17.5% against a full-data value of 16.2%, with the widest single-study swing 1.9 percentage points, so no one entry is carrying the result. Two pre-planned sensitivity runs tell the same story, since keeping only one row per publication gives +15.6% on six rows, and dropping the three entries whose intervals we derived rather than read gives +14.0% on nine rows. The pooled number is stable. Stability and meaning are different properties, and this pool has the first without the second.

The split that does mean something is by mechanism. Pooling the nine rearranging practices separately gives +10.9% (95% CI 7.6 to 14.4) with \(I^2 = 95.4\%\), while pooling the three importing practices gives +35.6% (95% CI 33.4 to 37.9) with \(I^2 = 0.0\%\) and a \(\tau^2\) that DerSimonian-Laird truncates to exactly zero. Three studies agreeing that closely is partly luck, but the contrast with the other nine is a hard one to look away from, and the difference between the groups runs to 0.20 log units with \(z = 11.3\).

Section 11 · the case for going deeper still

The Strongest Case Against Everything We Have Said

We asked one member to build the best available argument that this article is wrong, with permission to be as unfair to the rest of us as the evidence allowed, and what follows is that argument, parts of which we think land.

First, the depth critique proves less than it is made to prove. "No net gain to 40 centimetres" is a statement about a particular set of trials, most of them under 20 years old, in a soil profile that takes far longer than that to re-equilibrate below the plough layer. Luo and colleagues found a deep loss of 3.30 Mg ha⁻¹, but that loss is the old plough-enriched layer relaxing toward what it would have been without ploughing, and it is a one-time transition, not a permanent drain [7]. Once it has finished the surface gain continues and the deep loss does not, so a trial that happens to catch the middle of that transition records a null and then gets quoted as though it had found an absence.

Second, a null is not a zero. Haddaway and colleagues report no detectable effect across the full profile, but full-profile measurements are extremely noisy, since deep soil carbon varies enormously over short distances and the confidence intervals on whole-profile comparisons are wide enough to contain effects that would be commercially significant [8]; their own 0 to 30 centimetre estimate of 4.6 Mg ha⁻¹ has a 95% interval running from 0.78 to 8.43, and the full-profile interval is wider still. Failing to detect something is evidence of a small effect only when the study was capable of detecting a large one.

Third, we have quietly dismissed the ancillary benefits that actually matter to a farm. Ogle and colleagues, reviewing the same contested literature, conclude that no-till is better understood as a method for reducing erosion, adapting to a changing climate and protecting food security, with any carbon gain a co-benefit rather than the point [16, and see 13], and on that framing an article measuring no-till against a carbon yardstick and finding it wanting has measured the wrong thing very carefully indeed.

Fourth, our own pool undercuts our own conclusion. We pooled twelve effect sizes from six publications, four of them from a single second-order review, with two intervals we derived ourselves, and we report \(I^2 = 98\%\), which by our own account is not a credible synthesis. A review that says "the pooled number is nearly useless" and then draws a strong conclusion about mechanism from subgroup comparisons within that same pool is helping itself to precision it has just spent five sections denying.

Taking these in turn. The transition argument is the best of them and we cannot refute it with the data we have; it makes a testable prediction, that the deep loss should attenuate in trials running past thirty years, and somebody with access to the primary datasets should go and test it. The null-is-not-zero point is correct, and it is why we lean on the paired within-study contrasts rather than on whole-profile nulls, while the ancillary-benefits point we largely accept, as §12 says. The fourth objection, about our own inconsistency, is the one we want to answer directly: the subgroup contrast survives because a threefold difference in effect size, holding across depth strata, predicted in advance by a mechanism, with a control group behaving as that mechanism requires, is not the sort of thing noise produces. We would not defend the number 3.3, but we would defend the sign and the order of magnitude.

Section 12 · standing back from the pit

What We Think, and What We Would Pay For

No-till is good farming. Cover crops are good farming. Leaving residues on the field is good farming. Running a longer, more varied rotation is good farming. All four reduce erosion, improve water infiltration, feed soil biology and make a field more forgiving in a bad year, and every one of those benefits is realised on the farm by the person doing the work. None of that is in dispute here and none of it depends on a carbon market.

What we doubt is the specific claim that these practices reliably bank a measurable, additional, permanent tonnage of carbon that a third party can buy. Our reading of the evidence is that the tonnage is smaller than advertised, that a large part of the advertised figure comes from measuring only the depth where the carbon has been concentrated rather than the depth over which it has been redistributed, that the choice between fixed-depth and equivalent-soil-mass accounting can double or halve the answer in the same dataset [1, 22], and that whatever gain is real runs out on a timescale of decades because soils fill up [4, 5, 20, 21].

Powlson and colleagues said most of this in 2014 and were, as far as we can tell, correct [19]. Eleven years of further evidence has narrowed the range and not moved the centre.

Here is the part that bothers us most, and it has nothing to do with statistics. A carbon contract asks a farmer to accept a measurement clause they did not write, cannot audit and in most cases cannot afford to verify independently. Specify 30 centimetre fixed-depth cores and the farmer is paid for a number the soil science literature has known to be inflated since at least 2007; specify 60 centimetres and equivalent soil mass, which is what the people who study this actually recommend [22], and the measurement costs more while the payment shrinks, possibly to the point where the whole arrangement stops being worth anybody's time. Those are the two options, and no third one exists in which careful measurement and a large payment sit together.

What we would pay for, if anybody asked us, is more of what the meta-analyses are made of rather than more meta-analyses. Rothamsted has plots that have been under the same treatment since the 1840s, and Poulton and colleagues could say something definite about saturation only because somebody kept those plots going through two world wars and a great many funding rounds [20]. Perhaps a few dozen such experiments exist in the world, each costing a rounding error against the sums now moving through voluntary carbon markets, and each producing, slowly and undramatically, the only kind of evidence that can settle an argument about what a soil does over forty years.

Fund the long trials. Sample to 60 centimetres. Use equivalent soil mass. Then come back and tell us the number, and we will happily rewrite this article.

References

  1. Du, Z., Angers, D. A., Ren, T., Zhang, Q. & Li, G. (2017). The effect of no-till on organic C storage in Chinese soils should not be overemphasized: A meta-analysis. Agriculture, Ecosystems & Environment 236, 1–11. doi:10.1016/j.agee.2016.11.007
  2. Beillouin, D., Corbeels, M., Demenois, J., Berre, D., Boyer, A., Fallot, A., Feder, F. & Cardinael, R. (2023). A global meta-analysis of soil organic carbon in the Anthropocene. Nature Communications 14, 3700. doi:10.1038/s41467-023-39338-z
  3. Jian, J., Du, X., Reiter, M. S. & Stewart, R. D. (2020). A meta-analysis of global cropland soil carbon changes due to cover cropping. Soil Biology and Biochemistry 143, 107735. doi:10.1016/j.soilbio.2020.107735
  4. Liu, C., Lu, M., Cui, J., Li, B. & Fang, C. (2014). Effects of straw carbon input on carbon dynamics in agricultural soils: a meta-analysis. Global Change Biology 20, 1366–1381. doi:10.1111/gcb.12517
  5. Han, P., Zhang, W., Wang, G., Sun, W. & Huang, Y. (2016). Changes in soil organic carbon in croplands subjected to fertilizer management: a global meta-analysis. Scientific Reports 6, 27199. doi:10.1038/srep27199
  6. Gross, A. & Glaser, B. (2021). Meta-analysis on how manure application changes soil organic carbon storage. Scientific Reports 11, 5516. doi:10.1038/s41598-021-82739-7
  7. Luo, Z., Wang, E. & Sun, O. J. (2010). Can no-tillage stimulate carbon sequestration in agricultural soils? A meta-analysis of paired experiments. Agriculture, Ecosystems & Environment 139, 224–231. doi:10.1016/j.agee.2010.08.006
  8. Haddaway, N. R., Hedlund, K., Jackson, L. E., Kätterer, T., Lugato, E., Thomsen, I. K., Jørgensen, H. B. & Isberg, P.-E. (2017). How does tillage intensity affect soil organic carbon? A systematic review. Environmental Evidence 6, 30. doi:10.1186/s13750-017-0108-9
  9. Angers, D. A. & Eriksen-Hamel, N. S. (2008). Full-inversion tillage and organic carbon distribution in soil profiles: a meta-analysis. Soil Science Society of America Journal 72, 1370–1374. doi:10.2136/sssaj2007.0342
  10. Virto, I., Barré, P., Burlot, A. & Chenu, C. (2012). Carbon input differences as the main factor explaining the variability in soil organic C storage in no-tilled compared to inversion tilled agrosystems. Biogeochemistry 108, 17–26. doi:10.1007/s10533-011-9600-4
  11. McClelland, S. C., Paustian, K. & Schipanski, M. E. (2021). Management of cover crops in temperate climates influences soil organic carbon stocks: a meta-analysis. Ecological Applications 31, e02278. doi:10.1002/eap.2278
  12. Prairie, A. M., King, A. E. & Cotrufo, M. F. (2023). Restoring particulate and mineral-associated organic carbon through regenerative agriculture. Proceedings of the National Academy of Sciences 120, e2217481120. doi:10.1073/pnas.2217481120
  13. Bai, X., Huang, Y., Ren, W., Coyne, M., Jacinthe, P.-A., Tao, B., Hui, D., Yang, J. & Matocha, C. (2019). Responses of soil carbon sequestration to climate-smart agriculture practices: a meta-analysis. Global Change Biology 25, 2591–2606. doi:10.1111/gcb.14658
  14. Crystal-Ornelas, R., Thapa, R. & Tully, K. L. (2021). Soil organic carbon is affected by organic amendments, conservation tillage, and cover cropping in organic farming systems: a meta-analysis. Agriculture, Ecosystems & Environment 312, 107356. doi:10.1016/j.agee.2021.107356
  15. Emde, D., Hannam, K. D., Most, I., Nelson, L. M. & Jones, M. D. (2021). Soil organic carbon in irrigated agricultural systems: a meta-analysis. Global Change Biology 27, 3898–3910. doi:10.1111/gcb.15680
  16. Sun, W., Canadell, J. G., Yu, L., Yu, L., Zhang, W., Smith, P., Fischer, T. & Huang, Y. (2020). Climate drives global soil carbon sequestration and crop yield changes under conservation agriculture. Global Change Biology 26, 3325–3335. doi:10.1111/gcb.15001
  17. West, T. O. & Post, W. M. (2002). Soil organic carbon sequestration rates by tillage and crop rotation: a global data analysis. Soil Science Society of America Journal 66, 1930–1946. doi:10.2136/sssaj2002.1930
  18. Poeplau, C. & Don, A. (2015). Carbon sequestration in agricultural soils via cultivation of cover crops: a meta-analysis. Agriculture, Ecosystems & Environment 200, 33–41. doi:10.1016/j.agee.2014.10.024
  19. Powlson, D. S., Stirling, C. M., Jat, M. L., Gerard, B. G., Palm, C. A., Sanchez, P. A. & Cassman, K. G. (2014). Limited potential of no-till agriculture for climate change mitigation. Nature Climate Change 4, 678–683. doi:10.1038/nclimate2292
  20. Poulton, P., Johnston, J., Macdonald, A., White, R. & Powlson, D. (2018). Major limitations to achieving "4 per 1000" increases in soil organic carbon stock in temperate regions: evidence from long-term experiments at Rothamsted Research, United Kingdom. Global Change Biology 24, 2563–2584. doi:10.1111/gcb.14066
  21. Georgiou, K., Jackson, R. B., Vindušková, O., Abramoff, R. Z., Ahlström, A., Feng, W., Harden, J. W., Pellegrini, A. F. A., Polley, H. W., Soong, J. L., Riley, W. J. & Torn, M. S. (2022). Global stocks and capacity of mineral-associated soil organic carbon. Nature Communications 13, 3797. doi:10.1038/s41467-022-31540-9
  22. Raffeld, A. M., Bradford, M. A., Jackson, R. D., Rath, D., Sanford, G. R., Tautges, N. & Oldfield, E. E. (2024). The importance of accounting method and sampling depth to estimate changes in soil carbon stocks. Carbon Balance and Management 19, 2. doi:10.1186/s13021-024-00249-1