REVIEW · META-ANALYSIS · SOIL SCIENCE
How Much Carbon Can Farmland Actually Hold? A Review of the Field Trial Evidence
Review · Peer-edited by the club review board · LaTeX source · Our calculation · Interactive model
Section 1 · surface, 0–5 cm
The Number That Made Us Suspicious
Half a tonne per hectare per year.
Roughly the figure a farmer sees when a soil carbon programme comes calling, roughly the figure we came to check, and on its face nothing about it looks outrageous. Multiply it by the 44/12 conversion from carbon to carbon dioxide and you get about 1.8 tonnes of CO₂ equivalent per hectare per year, which at fifty dollars a tonne is ninety dollars a hectare, which on a decent-sized arable holding is real money arriving for doing something you were half-minded to do anyway. The pitch writes itself: stop ploughing, plant a cover crop over winter, leave the straw where it falls, drill straight into the residue next spring, and the soil quietly banks carbon that somebody else pays you for.
What made us suspicious was not the agronomy, which is sound and in places genuinely elegant. Our suspicion came from watching the same number surface again and again in documents whose authors had every commercial reason to want it large, and surface without the two things any measured quantity should carry with it: how deep they dug, and how long they waited.
So we went and looked, and found that the whole estimate turns on the depth question, which nobody had told us was anything more than a technicality. No-till does something real and well documented to a soil profile, concentrating organic matter near the surface, so that a survey sampling only near the surface measures a gain the whole profile does not have. Soil science has known this since at least 2007 [19]. The market has been slower to catch up.
This article is a meta-analysis, so the sources are the data: everything we did sits in the script, every number we extracted sits in §4 with a note saying which sentence of which paper it came out of, and §3 gives an honest account of why our pooled figure deserves less weight than the individual long-term experiments underneath it. Those experiments are the real heroes here, several of them running since the nineteenth century on funding nobody enjoys renewing, and they are the only reason any of this is arguable rather than merely assertable.
Section 2 · plough layer, 0–30 cm
What a Soil Carbon Credit Is, Mechanically
A credit is a promise about a quantity, so start with the quantity. Soil organic carbon, SOC, is the carbon held in dead and decaying plant and microbial material in the soil, as distinct from the carbonate minerals that also contain carbon but do not respond to farming on any timescale that matters here. Laboratories report it one of two ways: as a concentration, grams of carbon per kilogram of soil, which is what a combustion analyser actually measures, or as a stock, tonnes of carbon per hectare, which is the only one of the two that a market can sell.
Turning the first into the second requires the bulk density, the mass of a given volume of soil, and there the trouble starts, because a stock is concentration times bulk density times depth. Sample a 30 centimetre core, measure 12 g C per kg and a bulk density of 1.30 g/cm³, and you have 12 × 1.30 × 30 × 10 = 4,680 g C per square metre, or 46.8 tonnes per hectare. Clean arithmetic.
Now stop ploughing that field for fifteen years. Earthworms return and aggregates stabilise, the soil loosens, and the bulk density in the top layer falls to 1.20, so your 30 centimetre core now holds less soil than it used to. Hold the carbon concentration exactly constant and the calculated stock has still fallen; let the concentration rise a little and you may record a gain, a loss or nothing at all, depending entirely on which of the two changes won.
Soil scientists fixed this decades ago by comparing equivalent soil masses rather than equivalent depths: dig deeper in the loosened soil until the sample holds the same mass of mineral matter, then compare. Raffeld and colleagues, working on twenty years of data from two long-running American cropping trials, found that fixed-depth and equivalent-soil-mass estimates of stock change can differ by over a hundred per cent, and that the difference tracks the bulk-density change in the surface soil almost perfectly, with a correlation of 0.90 in one set of treatments [22]. They recommend equivalent soil mass, sampled to 60 cm in multiple increments. Most market protocols require neither.
Section 3 · no depth, this is about us
What We Did, and Everything That Is Wrong With It
We searched for published syntheses and multi-site analyses reporting soil organic carbon change under no-till, reduced tillage, cover cropping, crop residue or straw retention and crop rotation, and we deliberately also collected estimates for manure and for organic amendments substituted for mineral fertiliser, because those practices physically import carbon onto the field and so act as a control group for the central question. If a practice that genuinely adds carbon behaves one way with depth while a practice that only moves carbon around behaves another, the difference between them is evidence about mechanism rather than only about magnitude.
One person extracted each number and a second checked it against the source text, and every extracted value is quoted in §4 alongside the sentence it came out of.
Now the list of what is wrong with that, in descending order of how much it should worry you.
Our entries are not independent. This is the serious one. Published syntheses of the same literature share primary studies, sometimes heavily, and every pooling formula in the script assumes the effect sizes are independent draws, which ours plainly are not. Beillouin and colleagues explicitly report the fraction of primary studies used by more than one meta-analysis in their own dataset, which tells you they know the problem is real [2]. We report a pooled number because the brief asked for one and because the spread it exposes is informative. Nobody should price anything off it.
We screened abstracts rather than papers. A systematic review has a registered protocol, two people screening every record independently, a documented disagreement-resolution procedure and attempts to reach authors for unpublished data, and we have none of that. What we had was a search engine, an institutional-free open-access route to full texts when the publisher allowed it, and a Tuesday afternoon.
The discard pile is bigger than the pool, and it is not a random sample. We extracted 27 effect sizes and could pool 12, and what got discarded skews toward nulls, because a null is often reported as the bare sentence "no significant effect was detected" with no point estimate at all, and toward absolute stock differences in tonnes per hectare, which cannot be converted to a proportional change without a control stock the paper never gave. Both of those categories are systematically smaller than what survived, so our pooled estimate is biased upward by our own inclusion rule, and we cannot tell you by how much.
Two of our twelve intervals are derived rather than reported. Liu and colleagues give "12.8 ± 0.4%" without saying whether the ± is a standard error or a confidence interval, so we read it as a standard error and widened it [4]; Han and colleagues give an absolute change of 2.0 g/kg with an interval of 1.9 to 2.2 alongside a relative change of 19.5%, so we scaled the absolute interval by the ratio [5]. Both are flagged in the table, and both are dropped in a sensitivity run in §10.
A club extracting effect sizes from abstracts is not a systematic review, and we would rather say so in the third section than bury it in the twelfth.
Section 4 · the whole sampled column
The Ledger
Here is every number, where it came from and whether we could use it, the pool in the top block and the discard pile in the bottom one, the latter listed in full because a review that shows you only its survivors is showing you a selected sample and calling it a census.
| Key | Source | Practice | Effect | 95% interval | Reported n | Max depth | Provenance |
|---|---|---|---|---|---|---|---|
| Pooled · proportional change in soil organic carbon | |||||||
| Du17-ESM | Du et al. [1] | no-till, equivalent soil mass | +3.8% | 1.4 to 6.3 | 95 comparisons, 57 sites | 40 cm | abstract, verbatim |
| Du17-FD | Du et al. [1] | no-till, fixed depth | +5.1% | 2.5 to 7.7 | 95 comparisons, 57 sites | 30 cm | abstract, verbatim |
| Bei23-NT | Beillouin et al. [2] | no-till | +9.3% | 5.6 to 13.0 | 230 meta-analyses | not stated | results text, verbatim |
| Bei23-RT | Beillouin et al. [2] | reduced tillage | +12.0% | 0.1 to 24.0 | second-order | not stated | results text, verbatim |
| Bei23-RES | Beillouin et al. [2] | crop residue retention | +13.0% | 9.8 to 16.0 | second-order | not stated | results text, verbatim |
| Bei23-ROT | Beillouin et al. [2] | crop rotation | +6.5% | −0.9 to 14.0 | second-order | not stated | results text, verbatim |
| Jian20-CC | Jian et al. [3] | cover cropping | +15.5% | 13.8 to 17.3 | not stated in abstract | 30 cm | abstract, verbatim |
| Liu14-STR | Liu et al. [4] | straw return | +12.8% | 12.0 to 13.6 | 176 field studies | 20 cm | derived from ±0.4 |
| Han16-CFS | Han et al. [5] | straw plus fertiliser | +19.5% | 18.5 to 21.5 | global compilation | 20 cm | scaled from g/kg |
| Han16-CFM | Han et al. [5] | manure plus fertiliser | +36.2% | 33.1 to 39.3 | global compilation | 20 cm | scaled from g/kg |
| GG21-MAN | Gross & Glaser [6] | manure | +35.0% | 32.0 to 39.0 | 592 comparisons, 101 studies | 30 cm | results text, verbatim |
| Bei23-ORG | Beillouin et al. [2] | organic for mineral fertiliser | +34.0% | 20.0 to 49.0 | second-order | not stated | results text, verbatim |
| Pooled separately · annual accrual rate, Mg C ha⁻¹ yr⁻¹ | |||||||
| WP02 | West & Post [17] | no-till | 0.57 | ±0.14 (s.e.) | 276 paired treatments | mostly 30 cm | 57 ± 14 g C m⁻² yr⁻¹ |
| PD15 | Poeplau & Don [18] | cover crops | 0.32 | ±0.08 (s.e.) | 139 plots, 37 sites | 22 cm mean | abstract, verbatim |
| Du17-FDr | Du et al. [1] | no-till, fixed depth | 0.300 | 0.054 to 0.547 | 95 comparisons | 30 cm | abstract, verbatim |
| Du17-ESMr | Du et al. [1] | no-till, equivalent soil mass | 0.141 | −0.102 to 0.384 | 95 comparisons | 30 cm | abstract, verbatim |
| Extracted, read, and not poolable | |||||||
| Luo et al. [7] | no-till, 0–10 cm | +3.15 Mg/ha | ±2.42 | 69 paired experiments | 10 cm | stock, no control stock | |
| Luo et al. [7] | no-till, 20–40 cm | −3.30 Mg/ha | ±1.61 | 69 paired experiments | 40 cm | stock, no control stock | |
| Luo et al. [7] | no-till, 0–40 cm | no increase | none given | 69 paired experiments | 40 cm | reported as a null | |
| Haddaway et al. [8] | no-till, 0–30 cm, ≥10 yr | +4.6 Mg/ha | 0.78 to 8.43 | boreo-temperate | 30 cm | stock, no control stock | |
| Haddaway et al. [8] | no-till, full profile | no effect | none given | boreo-temperate | full | reported as a null | |
| Angers & Eriksen-Hamel [9] | no-till vs inversion tillage | +4.9 Mg/ha | not retrieved | paired, ≥5 yr, ≥30 cm | profile | no interval available | |
| Virto et al. [10] | no-till, 0–30 cm | +6.7% | not reported | 92 paired cases | 30 cm | no interval in abstract | |
| McClelland et al. [11] | cover crops, 0–30 cm | +12% | figure only | 181 obs, 40 papers | 30 cm | interval in a figure | |
| Prairie et al. [12] | no-till, 0–20 cm | +11.3% | figure only | 118 studies | 20 cm | interval in Fig. 2 | |
| Prairie et al. [12] | no-till, below 20 cm | not significant | none given | 118 studies | >20 cm | reported as a null | |
| Bai et al. [13] | conservation tillage | +5% | not reported | 3,049 paired measurements | mixed | no interval in abstract | |
| Bai et al. [13] | cover crops | +6% | not reported | 3,049 paired measurements | mixed | no interval in abstract | |
| Crystal-Ornelas et al. [14] | conservation tillage, organic | +14% | not reported | organic systems | weighted | no interval in abstract | |
| Emde et al. [15] | irrigation | +5.9% | figure only | 2–47 yr studies | mixed | interval in Fig. 3 | |
| Sun et al. [16] | no-till | regional only | none given | 260 paired studies | mixed | no global summary | |
Twenty-two sources read. Twenty-seven effect sizes extracted. Twelve poolable, from six publications. Those three numbers are the honest description of this review, and any reader who stops here has the important part.
Section 5 · 0–30 cm, where the studies are
Pooling, and Why the Pooled Number Should Not Be Used
We wrote the estimator from the formulas rather than calling a package, partly so that the club could see what the arithmetic actually does and partly because a routine you have not checked is a routine you have no business trusting. Effects are log response ratios, \(y = \ln(1 + p/100)\) for a percentage change \(p\), because response ratios multiply rather than add and their sampling distribution is far closer to normal in logs. Between-study variance is estimated by the DerSimonian-Laird moment estimator,
$$\hat\tau^2 = \max\left(0,\ \frac{Q - (k-1)}{\sum w_i - \sum w_i^2 / \sum w_i}\right), \qquad Q = \sum_i w_i (y_i - \hat\theta_{\text{FE}})^2$$and the random-effects weights are \(w_i^* = 1/(v_i + \hat\tau^2)\). Before running it on anything real we made it recover an answer we already knew: twelve synthetic studies drawn from a true effect of exactly +12% with a true \(\tau^2\) of 0.0025 came back at +12.09% with a 95% interval of 8.72 to 15.56, which contains the truth. Repeating the whole exercise 2,000 times, the mean estimated \(\tau^2\) was 0.002498 against a true 0.0025, while interval coverage came out at 91.5% rather than the nominal 95%, a shortfall that is the known behaviour of this estimator rather than a fault in our code, since it treats \(\hat\tau^2\) as though it were known exactly when it is nothing of the kind. Our intervals on the real data are therefore slightly too narrow, which we would rather print than hide.
Second check, and this one is an algebraic identity rather than a simulation. Hand the routine a dataset with no between-study heterogeneity at all and \(\hat\tau^2\) must come out at zero; at \(\hat\tau^2 = 0\) the random-effects weight \(1/(v_i + \tau^2)\) is the fixed-effect weight \(1/v_i\), so the two weighted means must be the same weighted mean. They are, to a difference of exactly 0.000e+00.
Now the real data, drawn as Figure 1.
Read the bottom bar, not the diamond: \(Q = 571.0\) on 11 degrees of freedom gives \(I^2 = 98.1\%\), which means essentially none of the spread between these studies is sampling noise. These are twelve measurements of twelve different quantities rather than twelve noisy measurements of one, taken in different soils, under different climates, over different durations and to different depths, so calling their weighted average "the effect of conservation agriculture on soil carbon" is a category error dressed up as a statistic.
The weights make the same point from a different angle: with \(\tau^2\) at 0.0051 and sampling variances two orders of magnitude smaller, \(\tau^2\) dominates every denominator and the weights flatten out to between 5.8% and 9.2%, so that a synthesis of 592 paired comparisons counts for about as much as one sentence lifted from a second-order review. Random-effects weighting behaves exactly that way when heterogeneity is enormous, the behaviour is correct, and it should still make you uncomfortable.
Section 6 · arguing, no fixed depth
Working Notes From the Club Table
R: if we include manure this stops being a review of carbon farming and becomes a review of everything.
K: manure is the control. If no-till and manure both show a topsoil gain and only manure keeps it at 40 cm, we've learned something about mechanism. Without manure we've got a list.
R: fine but label it. Don't let it sit in the same pool unmarked.
Agreed: two families, move and add, flagged in the data structure, pooled together once and separately once.
Four of our twelve rows come out of one paper [2]. K says that's double-counting and we should take one. R says it's a second-order review so its rows are genuinely different practices and the alternative is a pool of nine.
Compromise nobody loved: keep all four, then run a pre-planned one-row-per-publication sensitivity. It comes back at +15.6% on six rows against +16.2% on twelve. So the double-counting is not what is driving the answer. The heterogeneity is.
Session three · the sentence that changed the articleSomebody read the Du abstract out loud [1]. Same 95 comparisons. Same 57 sites. Fixed depth gives 0.300 Mg C ha⁻¹ yr⁻¹ and it is significant. Equivalent soil mass gives 0.141 and it is not.
K: so the difference between a credit and no credit is which arithmetic you use.
Long pause. Then twenty minutes of somebody insisting it must be a typo. The number is printed in the abstract, twice, and both figures carry their own confidence intervals.
We restructured the article around that sentence, which is now §8.
Session four · the thing we could not fixOnly seven of our twelve rows state a maximum sampling depth at all. Seven. In a literature whose central methodological dispute is about sampling depth.
R: that's not a limitation of our review, that's a finding about the field.
We wrote it up as one, and it opens §7.
Session five · on being annoying about itNote from the review board, verbatim: "You are three drafts into sounding like people who think farmers are being fooled. They are not being fooled. They are being offered a contract with a measurement clause they did not write and cannot audit. Fix the tone."
Fixed, we think. The scepticism belongs to the accounting, not to anybody standing in a field.
Section 7 · 30–60 cm, below the plough
The Depth Problem
Seven of our twelve pooled effect sizes state a maximum sampling depth, and the other five do not state one anywhere we could find. The first result of this section took us by surprise, and it is worse than it sounds, because a synthesis that does not report its depth range has usually not restricted it, which in practice means it is dominated by studies that sampled the top 20 or 30 centimetres, that being what most trials did.
Here is what happens to a soil profile under no-till: residues stay on the surface instead of being turned under, so organic matter accumulates in the top few centimetres and the profile becomes strongly stratified, while the layer the plough used to enrich, the bottom of the old furrow slice around 20 to 30 centimetres, stops receiving buried residue and loses carbon. Luo and colleagues measured exactly this across 69 paired experiments where sampling went below 40 centimetres: a gain of 3.15 ± 2.42 Mg C ha⁻¹ in the surface 10 centimetres, a loss of 3.30 ± 1.61 Mg C ha⁻¹ in the 20 to 40 centimetre layer, and no increase in total carbon down to 40 centimetres [7]. Figure 2 is that result at scale.
That pattern is not one team's result; it recurs whenever anybody looks. Prairie, King and Cotrufo, synthesising 118 studies and 157 experiments, found no-till raised topsoil carbon by 11.3% from 0 to 20 centimetres and had no significant effect below 20 [12]. Jian and colleagues found cover cropping significantly raised carbon in soils at or above 30 centimetres and not below [3], while Haddaway and colleagues found 4.6 Mg ha⁻¹ in the top 30 centimetres over comparisons of at least ten years and no effect whatever across the full profile [8]. Angers and Eriksen-Hamel, restricting themselves to replicated randomised paired trials of at least five years measured to at least 30 centimetres, found the surface gain under no-till was largely offset by greater carbon near the bottom of the plough layer under inversion tillage [9].
Now the control. Gross and Glaser pooled 592 paired comparisons of manure application, which is carbon physically carted onto the field from somewhere else, and found a 40% increase where sampling stopped at 15 centimetres against a 23% increase below 30 centimetres [6]. Smaller at depth, certainly, but present, significant, and not reversing sign.
Five practices that move carbon lose the effect below the plough layer. One practice that brings carbon in keeps it. The whole argument fits in two sentences.
We tried to test this as a meta-regression of effect size on maximum sampling depth across the seven rows that state one, and the slope comes out at −0.0071 in log units per centimetre, which works out at −6.9% per additional ten centimetres of required depth, with \(p = 0.20\). We are not going to dress that up. With seven points and exactly one beyond 30 centimetres the regression is underpowered, and the honest statement is that the between-study evidence is too thin to test the thing the field most needs tested, which is why the within-study paired contrasts in the paragraphs above carry far more weight here: they hold soil, climate, crop and duration constant and vary only depth.
One calculation makes the stakes visible, and Figure 3 draws it. Suppose a credit protocol refused to count any study that had not sampled to at least a given depth: what survives, and what does the remainder say?
That last point is the article in one mark on a chart. Demand that a study sample to 40 centimetres and correct for bulk density, and our entire twelve-entry dataset reduces to one number, which is four times smaller than the pooled estimate and comes with the observation, from the same authors in the same abstract, that the effect on the annual rate is not statistically distinguishable from zero [1].
Section 8 · 0–30 cm, the core that was taken
Plain Arithmetic
This section states numbers. Du and colleagues analysed 95 paired comparisons from 57 experimental sites in China, all running three years or more [1]. They computed the no-till effect on the top 30 centimetres two ways.
rate 0.300 Mg C ha⁻¹ yr⁻¹ 95% CI 0.054 to 0.547 p < 0.001
equivalent soil mass
rate 0.141 Mg C ha⁻¹ yr⁻¹ 95% CI −0.102 to 0.384 p = 0.100
ratio 0.300 / 0.141 = 2.13
difference 0.300 − 0.141 = 0.159 Mg C ha⁻¹ yr⁻¹
Same fields. Same cores. Same laboratory. The second figure corrects for no-till soil being less dense, so a fixed 30 centimetre core holds less soil under no-till than under the plough.
Carbon to carbon dioxide is 44/12 = 3.667.
0.141 × 3.667 = 0.517 t CO₂e ha⁻¹ yr⁻¹
at $50 per t CO₂e
fixed depth $55.00 per hectare per year
equivalent soil mass $25.85 per hectare per year
200 hectares × $25.85 = $5,170 per year, gross of sampling, laboratory, verification and registry fees
For comparison, the two rate estimates most often cited. West and Post, from 67 long-term experiments and 276 paired treatments, give 57 ± 14 g C m⁻² yr⁻¹ for a change from conventional tillage to no-till, which is 0.57 ± 0.14 Mg C ha⁻¹ yr⁻¹ [17]. Poeplau and Don, from 139 plots at 37 sites, give 0.32 ± 0.08 Mg C ha⁻¹ yr⁻¹ for cover crops at a mean sampling depth of 22 centimetres [18].
0.324 Mg C ha⁻¹ yr⁻¹ 95% CI 0.175 to 0.473
Q = 5.30 on 3 df, τ² = 0.0100, I² = 43.4%
= 1.186 t CO₂e ha⁻¹ yr⁻¹ = $59.32 per hectare per year at $50
The heterogeneity in this small pool is 43.4%, which is moderate rather than catastrophic, and the reason is that all four estimates come from the same shallow depth band. Restrict the question enough and the studies start agreeing. Restricting the question is not the same as answering it.
Section 9 · whole profile, over decades
The Ceiling
Even if every number above were larger, there would be a second problem, and it is the one that long-term experiments are uniquely placed to see. Soil carbon does not accumulate forever.
Write the simplest model anybody could write: carbon enters the stable soil pool at some annual rate \(I\) set by plant residues and roots, and leaves through microbial respiration at a rate proportional to how much is already there, with fractional loss constant \(k\). Then
$$\frac{dC}{dt} = I - kC, \qquad C(t) = C_{\text{eq}} + (C_0 - C_{\text{eq}})e^{-kt}, \qquad C_{\text{eq}} = \frac{I}{k}$$Raising the input raises the equilibrium in exact proportion, and the approach to that equilibrium is exponential. The practice does not stop working because the farmer stops trying; it stops working because the stock has climbed close enough to its new ceiling that respiration has caught up with the extra input, and by then there is nothing left to sell. Figure 4 draws one soil doing it.
Nothing in that model was fitted to anything: it carries two constants and one line of calculus, and the published durations line up with it uncomfortably well. Liu and colleagues, from 176 field studies, put saturation under straw return at around twelve years [4], and Han and colleagues estimate sequestration durations of 28 to 73 years for straw plus fertiliser and 26 to 117 years for manure plus fertiliser, with wide variation across climates [5]. Poeplau and Don, running their cover-crop data through a turnover model, get a new steady state after roughly 155 years with a total accumulation of 16.7 ± 1.5 Mg C ha⁻¹ at 22 centimetres depth, within a rounding error of the ceiling our two-constant model produces by accident [18].
The strongest evidence on ceilings comes from the experiments nobody wants to fund. Poulton and colleagues went through 114 treatment comparisons across 16 long-term experiments at Rothamsted, some running for 157 years, and asked whether the "4 per 1000" target of a 4 per mille annual increase in soil carbon was achievable [20]. In the two longest experiments, annual farmyard manure at 35 tonnes fresh weight per hectare produced increases of 18 and 43 per mille per year over the first twenty years, and rates above 7 per mille continued for another forty to sixty. Then they fell away. A soil receiving an enormous and unbroken carbon input from outside still slows down, because it is filling.
Georgiou and colleagues put a mechanism under this from 1,144 soil profiles worldwide: carbon bound to mineral surfaces is the durable fraction, mineral surfaces are finite, and soils currently sit at about 42% of their mineralogical capacity in surface layers and 21% at depth [21]. Across 103 carbon-accrual measurements they found that soils furthest from capacity accrue fastest, with rates averaging three times higher at a tenth of capacity than at a half, while our one-line model, given a 100 Mg ha⁻¹ capacity, produces a ratio of 4.4 for the same comparison. Right sign, roughly right size, out of a model told nothing whatever about minerals, which we take as encouragement rather than as confirmation.
Section 10 · turning the instrument on ourselves
Our Own Funnel, and What It Cannot Tell Us
A funnel plot puts each study's effect against its precision, so that precise studies cluster near the pooled value at the top while imprecise ones scatter symmetrically below, and the shape should be a funnel. When the bottom left of that funnel comes out empty, the usual reading is that small studies finding small or negative effects were run and never published. Figure 5 is ours.
Our funnel does not fail that test so much as decline to take it: the points form two clusters at different effect sizes rather than a funnel, because the dataset contains two mechanistically different kinds of practice. Egger's regression of the standard normal deviate on precision gives an intercept of +1.61 against a standard error of 3.56, so \(t = 0.45\) on 10 degrees of freedom and \(p = 0.66\). With twelve points and \(I^2\) at 98% that test has almost no power, and heterogeneity manufactures funnel asymmetry on its own even where no publication bias exists, so we report the number because we said we would, and nobody should read it as evidence of absence.
Leave-one-out is more informative: removing each study in turn moves the pooled estimate over a range of 14.3% to 17.5% against a full-data value of 16.2%, with the widest single-study swing 1.9 percentage points, so no one entry is carrying the result. Two pre-planned sensitivity runs tell the same story, since keeping only one row per publication gives +15.6% on six rows, and dropping the three entries whose intervals we derived rather than read gives +14.0% on nine rows. The pooled number is stable. Stability and meaning are different properties, and this pool has the first without the second.
The split that does mean something is by mechanism. Pooling the nine rearranging practices separately gives +10.9% (95% CI 7.6 to 14.4) with \(I^2 = 95.4\%\), while pooling the three importing practices gives +35.6% (95% CI 33.4 to 37.9) with \(I^2 = 0.0\%\) and a \(\tau^2\) that DerSimonian-Laird truncates to exactly zero. Three studies agreeing that closely is partly luck, but the contrast with the other nine is a hard one to look away from, and the difference between the groups runs to 0.20 log units with \(z = 11.3\).
Section 11 · the case for going deeper still
The Strongest Case Against Everything We Have Said
We asked one member to build the best available argument that this article is wrong, with permission to be as unfair to the rest of us as the evidence allowed, and what follows is that argument, parts of which we think land.
First, the depth critique proves less than it is made to prove. "No net gain to 40 centimetres" is a statement about a particular set of trials, most of them under 20 years old, in a soil profile that takes far longer than that to re-equilibrate below the plough layer. Luo and colleagues found a deep loss of 3.30 Mg ha⁻¹, but that loss is the old plough-enriched layer relaxing toward what it would have been without ploughing, and it is a one-time transition, not a permanent drain [7]. Once it has finished the surface gain continues and the deep loss does not, so a trial that happens to catch the middle of that transition records a null and then gets quoted as though it had found an absence.
Second, a null is not a zero. Haddaway and colleagues report no detectable effect across the full profile, but full-profile measurements are extremely noisy, since deep soil carbon varies enormously over short distances and the confidence intervals on whole-profile comparisons are wide enough to contain effects that would be commercially significant [8]; their own 0 to 30 centimetre estimate of 4.6 Mg ha⁻¹ has a 95% interval running from 0.78 to 8.43, and the full-profile interval is wider still. Failing to detect something is evidence of a small effect only when the study was capable of detecting a large one.
Third, we have quietly dismissed the ancillary benefits that actually matter to a farm. Ogle and colleagues, reviewing the same contested literature, conclude that no-till is better understood as a method for reducing erosion, adapting to a changing climate and protecting food security, with any carbon gain a co-benefit rather than the point [16, and see 13], and on that framing an article measuring no-till against a carbon yardstick and finding it wanting has measured the wrong thing very carefully indeed.
Fourth, our own pool undercuts our own conclusion. We pooled twelve effect sizes from six publications, four of them from a single second-order review, with two intervals we derived ourselves, and we report \(I^2 = 98\%\), which by our own account is not a credible synthesis. A review that says "the pooled number is nearly useless" and then draws a strong conclusion about mechanism from subgroup comparisons within that same pool is helping itself to precision it has just spent five sections denying.
Taking these in turn. The transition argument is the best of them and we cannot refute it with the data we have; it makes a testable prediction, that the deep loss should attenuate in trials running past thirty years, and somebody with access to the primary datasets should go and test it. The null-is-not-zero point is correct, and it is why we lean on the paired within-study contrasts rather than on whole-profile nulls, while the ancillary-benefits point we largely accept, as §12 says. The fourth objection, about our own inconsistency, is the one we want to answer directly: the subgroup contrast survives because a threefold difference in effect size, holding across depth strata, predicted in advance by a mechanism, with a control group behaving as that mechanism requires, is not the sort of thing noise produces. We would not defend the number 3.3, but we would defend the sign and the order of magnitude.
Section 12 · standing back from the pit
What We Think, and What We Would Pay For
No-till is good farming. Cover crops are good farming. Leaving residues on the field is good farming. Running a longer, more varied rotation is good farming. All four reduce erosion, improve water infiltration, feed soil biology and make a field more forgiving in a bad year, and every one of those benefits is realised on the farm by the person doing the work. None of that is in dispute here and none of it depends on a carbon market.
What we doubt is the specific claim that these practices reliably bank a measurable, additional, permanent tonnage of carbon that a third party can buy. Our reading of the evidence is that the tonnage is smaller than advertised, that a large part of the advertised figure comes from measuring only the depth where the carbon has been concentrated rather than the depth over which it has been redistributed, that the choice between fixed-depth and equivalent-soil-mass accounting can double or halve the answer in the same dataset [1, 22], and that whatever gain is real runs out on a timescale of decades because soils fill up [4, 5, 20, 21].
Powlson and colleagues said most of this in 2014 and were, as far as we can tell, correct [19]. Eleven years of further evidence has narrowed the range and not moved the centre.
Here is the part that bothers us most, and it has nothing to do with statistics. A carbon contract asks a farmer to accept a measurement clause they did not write, cannot audit and in most cases cannot afford to verify independently. Specify 30 centimetre fixed-depth cores and the farmer is paid for a number the soil science literature has known to be inflated since at least 2007; specify 60 centimetres and equivalent soil mass, which is what the people who study this actually recommend [22], and the measurement costs more while the payment shrinks, possibly to the point where the whole arrangement stops being worth anybody's time. Those are the two options, and no third one exists in which careful measurement and a large payment sit together.
What we would pay for, if anybody asked us, is more of what the meta-analyses are made of rather than more meta-analyses. Rothamsted has plots that have been under the same treatment since the 1840s, and Poulton and colleagues could say something definite about saturation only because somebody kept those plots going through two world wars and a great many funding rounds [20]. Perhaps a few dozen such experiments exist in the world, each costing a rounding error against the sums now moving through voluntary carbon markets, and each producing, slowly and undramatically, the only kind of evidence that can settle an argument about what a soil does over forty years.
Fund the long trials. Sample to 60 centimetres. Use equivalent soil mass. Then come back and tell us the number, and we will happily rewrite this article.
References
- Du, Z., Angers, D. A., Ren, T., Zhang, Q. & Li, G. (2017). The effect of no-till on organic C storage in Chinese soils should not be overemphasized: A meta-analysis. Agriculture, Ecosystems & Environment 236, 1–11. doi:10.1016/j.agee.2016.11.007
- Beillouin, D., Corbeels, M., Demenois, J., Berre, D., Boyer, A., Fallot, A., Feder, F. & Cardinael, R. (2023). A global meta-analysis of soil organic carbon in the Anthropocene. Nature Communications 14, 3700. doi:10.1038/s41467-023-39338-z
- Jian, J., Du, X., Reiter, M. S. & Stewart, R. D. (2020). A meta-analysis of global cropland soil carbon changes due to cover cropping. Soil Biology and Biochemistry 143, 107735. doi:10.1016/j.soilbio.2020.107735
- Liu, C., Lu, M., Cui, J., Li, B. & Fang, C. (2014). Effects of straw carbon input on carbon dynamics in agricultural soils: a meta-analysis. Global Change Biology 20, 1366–1381. doi:10.1111/gcb.12517
- Han, P., Zhang, W., Wang, G., Sun, W. & Huang, Y. (2016). Changes in soil organic carbon in croplands subjected to fertilizer management: a global meta-analysis. Scientific Reports 6, 27199. doi:10.1038/srep27199
- Gross, A. & Glaser, B. (2021). Meta-analysis on how manure application changes soil organic carbon storage. Scientific Reports 11, 5516. doi:10.1038/s41598-021-82739-7
- Luo, Z., Wang, E. & Sun, O. J. (2010). Can no-tillage stimulate carbon sequestration in agricultural soils? A meta-analysis of paired experiments. Agriculture, Ecosystems & Environment 139, 224–231. doi:10.1016/j.agee.2010.08.006
- Haddaway, N. R., Hedlund, K., Jackson, L. E., Kätterer, T., Lugato, E., Thomsen, I. K., Jørgensen, H. B. & Isberg, P.-E. (2017). How does tillage intensity affect soil organic carbon? A systematic review. Environmental Evidence 6, 30. doi:10.1186/s13750-017-0108-9
- Angers, D. A. & Eriksen-Hamel, N. S. (2008). Full-inversion tillage and organic carbon distribution in soil profiles: a meta-analysis. Soil Science Society of America Journal 72, 1370–1374. doi:10.2136/sssaj2007.0342
- Virto, I., Barré, P., Burlot, A. & Chenu, C. (2012). Carbon input differences as the main factor explaining the variability in soil organic C storage in no-tilled compared to inversion tilled agrosystems. Biogeochemistry 108, 17–26. doi:10.1007/s10533-011-9600-4
- McClelland, S. C., Paustian, K. & Schipanski, M. E. (2021). Management of cover crops in temperate climates influences soil organic carbon stocks: a meta-analysis. Ecological Applications 31, e02278. doi:10.1002/eap.2278
- Prairie, A. M., King, A. E. & Cotrufo, M. F. (2023). Restoring particulate and mineral-associated organic carbon through regenerative agriculture. Proceedings of the National Academy of Sciences 120, e2217481120. doi:10.1073/pnas.2217481120
- Bai, X., Huang, Y., Ren, W., Coyne, M., Jacinthe, P.-A., Tao, B., Hui, D., Yang, J. & Matocha, C. (2019). Responses of soil carbon sequestration to climate-smart agriculture practices: a meta-analysis. Global Change Biology 25, 2591–2606. doi:10.1111/gcb.14658
- Crystal-Ornelas, R., Thapa, R. & Tully, K. L. (2021). Soil organic carbon is affected by organic amendments, conservation tillage, and cover cropping in organic farming systems: a meta-analysis. Agriculture, Ecosystems & Environment 312, 107356. doi:10.1016/j.agee.2021.107356
- Emde, D., Hannam, K. D., Most, I., Nelson, L. M. & Jones, M. D. (2021). Soil organic carbon in irrigated agricultural systems: a meta-analysis. Global Change Biology 27, 3898–3910. doi:10.1111/gcb.15680
- Sun, W., Canadell, J. G., Yu, L., Yu, L., Zhang, W., Smith, P., Fischer, T. & Huang, Y. (2020). Climate drives global soil carbon sequestration and crop yield changes under conservation agriculture. Global Change Biology 26, 3325–3335. doi:10.1111/gcb.15001
- West, T. O. & Post, W. M. (2002). Soil organic carbon sequestration rates by tillage and crop rotation: a global data analysis. Soil Science Society of America Journal 66, 1930–1946. doi:10.2136/sssaj2002.1930
- Poeplau, C. & Don, A. (2015). Carbon sequestration in agricultural soils via cultivation of cover crops: a meta-analysis. Agriculture, Ecosystems & Environment 200, 33–41. doi:10.1016/j.agee.2014.10.024
- Powlson, D. S., Stirling, C. M., Jat, M. L., Gerard, B. G., Palm, C. A., Sanchez, P. A. & Cassman, K. G. (2014). Limited potential of no-till agriculture for climate change mitigation. Nature Climate Change 4, 678–683. doi:10.1038/nclimate2292
- Poulton, P., Johnston, J., Macdonald, A., White, R. & Powlson, D. (2018). Major limitations to achieving "4 per 1000" increases in soil organic carbon stock in temperate regions: evidence from long-term experiments at Rothamsted Research, United Kingdom. Global Change Biology 24, 2563–2584. doi:10.1111/gcb.14066
- Georgiou, K., Jackson, R. B., Vindušková, O., Abramoff, R. Z., Ahlström, A., Feng, W., Harden, J. W., Pellegrini, A. F. A., Polley, H. W., Soong, J. L., Riley, W. J. & Torn, M. S. (2022). Global stocks and capacity of mineral-associated soil organic carbon. Nature Communications 13, 3797. doi:10.1038/s41467-022-31540-9
- Raffeld, A. M., Bradford, M. A., Jackson, R. D., Rath, D., Sanford, G. R., Tautges, N. & Oldfield, E. E. (2024). The importance of accounting method and sampling depth to estimate changes in soil carbon stocks. Carbon Balance and Management 19, 2. doi:10.1186/s13021-024-00249-1