REVIEW · META-ANALYSIS · ECOLOGY
The Insect Decline Literature Disagrees With Itself, and We Wanted to Know How Badly
Review · Peer-edited by the club review board · LaTeX source · Our calculation · Interactive model
Two Numbers That Cannot Both Be The Headline
The problem, in one paragraph.
In October 2017 a team working with amateur entomologists in Krefeld published Malaise-trap biomass from 63 German nature reserves and reported a 76.7 percent fall over 27 years, with a mid-summer figure of 81.6 percent [1]. A Malaise trap is a tent of fine mesh that flying insects blunder into and then crawl upwards into a collecting jar, and it does not care in the least what it catches. That indifference is the instrument's one virtue. The number went around the world in a day. In April 2020 a different team compiled 166 long-term surveys across 1,676 sites and reported a terrestrial decline of about 0.92 percent per year, roughly 9 percent per decade, with freshwater insects moving the other way at about +11 percent per decade [2]. Four months later, in August 2020, a third team pulled more than 5,300 arthropod time series out of the United States Long Term Ecological Research network and found net trends that were, in their words, generally indistinguishable from zero [6].
None of those three papers is bad. We read all of them, twice in two cases, and found the reasoning in each one about as careful as the reasoning in the other two. We would sign off on all of them.
And yet the first says one thing, the second says a smaller version of that thing in one place and the opposite of it in another, and the third, drawing on an archive that had kept the raw counts for decades, says nothing much is happening at all. The public conversation settled the matter the way it settles most matters of this shape, by picking whichever number suited the argument already under way. Press releases got the 76 percent. Columnists who dislike environmental alarm got the zero. Almost nobody quoted the middle estimate, the one built from the most data, because a 9 percent decline per decade makes a footnote rather than a headline.
We got tired of this. So we did the arithmetic ourselves.
What follows is a meta-analysis, which is to say a study whose data are other studies, carrying every error those studies carried plus a few of its own. We put published trend estimates onto a single scale, pooled them under a model that lets the true effect differ between sites, and then spent most of our effort on the diagnostics that say whether the pooling meant anything at all. Figure 1 shows why the disagreement is not cosmetic. Five published annual rates, run forward for thirty years, land anywhere between an 85 percent loss and a 38 percent gain, which is not a disagreement about magnitude but a disagreement about direction.
What A School Club Can And Cannot Do To A Literature
Before anyone quotes this, we should be exact about what it is, because a school club's arithmetic and a systematic review are different objects, and only one of them is on offer here.
A systematic review has a pre-registered search string, two people screening every title independently, a documented reason for every exclusion, and usually a request to the original authors for data that was never printed. We have none of that. Our search procedure was to chase reference lists until they repeated. Our screening procedure was six people at a table arguing, which is a worse procedure than two independent screeners and a better one than nobody noticing. Our data extraction was one person reading a table aloud while a second person typed, then the two swapping roles and doing the whole thing again to see whether the numbers came out the same.
That gives us a convenience sample. That sample is this article's worst flaw. No amount of correct arithmetic downstream repairs it. If the papers that came to mind first are the papers that made news, and papers make news by finding large effects, then the pool is tilted before we start. We test for that in §8. We do not entirely like the answer.
Against that, one rule we did keep, and kept strictly. Every effect size in the pool traces to a printed number in its source. Not read off a figure with a ruler, and not imputed from a percentage plus a guessed duration, which is the shortcut that would have let three more papers into the drawer. A study got into the pool if and only if it printed a log-linear annual trend coefficient with a standard error, or printed a percent-per-year with a 95 percent interval from which a standard error could be recovered by arithmetic that any reader can repeat. The provenance column in Table 1 quotes the sentence that carries the number, or names the table and the cell the number sits in, so a reader can find it without asking us anything. Check every row against its source in an afternoon. Tell us if we misread one.
We fixed the second rule before looking at any numbers, which matters more than it sounds, because a rule written afterwards will quietly select the effects that make the story tidier and nobody in the room will notice themselves doing it. One effect per independent monitoring dataset, always the paper's own headline series. We took whichever order the paper itself led with. Never the largest decline on offer. Two papers still contribute more than one effect each, because they hold genuinely separate monitoring programmes at separate sites, and §9 re-runs the whole pool with those collapsed to one apiece.
Third, and worst: we are not measuring insects. We are measuring published measurements of insects, a different quantity with an error structure of its own, and that difference widens until it becomes the whole subject by §11.
The Drawer
Eleven specimens. Ten localities, six papers, four countries, two metrics. Forty-nine years of calendar time separate the earliest start from the latest finish. Pinned and labelled below. A meta-analysis dataset is a collection of things other people caught, mounted by other people again, and a drawer is the honest way to display one.
Hallmann et al. 2017
- Locality
- 63 protected areas, lowland Germany
- Duration
- 1989–2016, 27 years, 1,503 trap samples
- Taxa
- all flying insects, Malaise-trap dry biomass
- Trend
- ρ = −0.0630 ± 0.0020 · −6.11 %/yr
Shortall et al. 2009 · Hereford
- Locality
- Hereford, western England
- Duration
- 1973–2002, 30 annual indices
- Taxa
- aerial insects, 12.2 m suction trap, biomass
- Trend
- ρ = −0.0434 ± 0.0094 · −4.25 %/yr
Shortall et al. 2009 · Rothamsted
- Locality
- Harpenden, Hertfordshire
- Duration
- 1973–2002, 30 annual indices
- Taxa
- aerial insects, 12.2 m suction trap, biomass
- Trend
- ρ = +0.0128 ± 0.0094 · +1.29 %/yr
Shortall et al. 2009 · Starcross
- Locality
- Starcross, Devon
- Duration
- 1973–2002, 30 annual indices
- Taxa
- aerial insects, 12.2 m suction trap, biomass
- Trend
- ρ = −0.0050 ± 0.0061 · −0.50 %/yr
Shortall et al. 2009 · Wye
- Locality
- Wye, Kent
- Duration
- 1973–2002, 30 annual indices
- Taxa
- aerial insects, 12.2 m suction trap, biomass
- Trend
- ρ = −0.0030 ± 0.0061 · −0.30 %/yr
Hallmann et al. 2020 · Kaaistoep
- Locality
- De Kaaistoep, North Brabant, Netherlands
- Duration
- 1997–2017, 497 trapping nights
- Taxa
- macro-moths at light, 54,492 individuals
- Trend
- ρ = −0.0400 ± 0.0060 · −3.92 %/yr
Hallmann et al. 2020 · Kaaistoep
- Locality
- De Kaaistoep, North Brabant, Netherlands
- Duration
- 1997–2017, 572 trapping nights
- Taxa
- beetles at light, 257,793 individuals
- Trend
- ρ = −0.0480 ± 0.0100 · −4.69 %/yr
Hallmann et al. 2020 · Wijster
- Locality
- Wijster, Drenthe, Netherlands
- Duration
- 1986–2016, 26 sampled years, 48 pitfalls
- Taxa
- ground beetles, 264,986 individuals, 156 species
- Trend
- ρ = −0.0440 ± 0.0060 · −4.30 %/yr
Wepprich et al. 2019
- Locality
- Ohio, USA, 104 monitored sites
- Duration
- 1996–2016, 24,405 volunteer surveys
- Taxa
- butterflies, 81 species with estimable trends
- Trend
- ρ = −0.0200 ± 0.0050 · −1.98 %/yr
Edwards et al. 2025
- Locality
- contiguous USA, 2,478 locations, 35 programmes
- Duration
- 2000–2020, 76,957 surveys
- Taxa
- butterflies, 554 species, 12.6 M individuals
- Trend
- ρ = −0.0131 ± 0.0054 · −1.30 %/yr
Weiss et al. 2024
- Locality
- old beech forest, north-east Germany
- Duration
- 1999–2022, 24-year pitfall series
- Taxa
- carabid beetles, total abundance
- Trend
- ρ = −0.0315 ± 0.0113 · −3.10 %/yr
Look at the four Shortall labels [8]. Same country, same instrument, same protocol, same three decades, same team doing the counting, and four answers that a reader would take for four different continents. Hereford loses biomass at 4.25 percent a year. Rothamsted gains it at 1.29 percent a year. Starcross and Wye do essentially nothing. Anyone wanting a single demonstration that the insect trend literature measures something spatially wild rather than a planetary constant can stop reading here, and might note that the demonstration was printed eight years before the Krefeld paper that started the argument.
Table 1 is the extraction sheet, and the only part of this article that would survive our conclusions being wrong. Everything above depends on it, so it gets printed here in full rather than deposited in a repository nobody opens, and the lower block lists the work we read, cited and deliberately did not pool, with the reason alongside each row.
| Study | Locality | Span | Taxa / metric | Sample | ρ yr−1 | SE | %/yr | Where the number is printed |
|---|---|---|---|---|---|---|---|---|
| Pooled · slope with standard error, or rate with a 95% interval | ||||||||
| Hallmann et al. 2017 [1] | 63 reserves, Germany | 1989–2016 | flying insects, Malaise biomass | 1,503 samples | −0.06300 | 0.00200 | −6.11 | Results: “annual trend coefficient = −0.063, sd = 0.002, i.e. 6.1% annual decline” |
| Shortall et al. 2009 [8] | Hereford, England | 1973–2002 | aerial biomass, suction trap | 30 indices | −0.04340 | 0.00944 | −4.25 | Table 1, row “Biomass Hereford”: slope −0.01885, SE 0.00410 (log10), × ln 10 |
| Shortall et al. 2009 [8] | Harpenden, England | 1973–2002 | aerial biomass, suction trap | 30 indices | +0.01283 | 0.00942 | +1.29 | Table 1, row “Biomass Rothamsted”: slope +0.00557, SE 0.00409 (log10) |
| Shortall et al. 2009 [8] | Starcross, England | 1973–2002 | aerial biomass, suction trap | 30 indices | −0.00500 | 0.00610 | −0.50 | Table 1, row “Biomass Starcross”: slope −0.00217, SE 0.00265 (log10) |
| Shortall et al. 2009 [8] | Wye, England | 1973–2002 | aerial biomass, suction trap | 30 indices | −0.00297 | 0.00610 | −0.30 | Table 1, row “Biomass Wye”: slope −0.00129, SE 0.00265 (log10) |
| Hallmann et al. 2020 [9] | De Kaaistoep, NL | 1997–2017 | macro-moths at light | 497 nights | −0.04000 | 0.00600 | −3.92 | Table 2, Lepidoptera: ρ −0.040 (0.006) |
| Hallmann et al. 2020 [9] | De Kaaistoep, NL | 1997–2017 | beetles at light | 572 nights | −0.04800 | 0.01000 | −4.69 | Table 2, Coleoptera: ρ −0.048 (0.010) |
| Hallmann et al. 2020 [9] | Wijster, NL | 1986–2016 | ground beetles, 48 pitfalls | 264,986 ind. | −0.04400 | 0.00600 | −4.30 | Results: “ρ = −0.044, se = 0.006, P < 0.001, 4% decline per year” |
| Wepprich et al. 2019 [10] | Ohio, USA | 1996–2016 | butterflies, 81 species | 24,405 surveys | −0.02000 | 0.00500 | −1.98 | Results: “declined at an annual rate of 2.0% (β1 = −0.020, std. err. 0.005, p < 0.001)” |
| Edwards et al. 2025 [11] | contiguous USA | 2000–2020 | butterflies, 554 species | 76,957 surveys | −0.01309 | 0.00543 | −1.30 | “a rate of 1.3% annually [95% confidence interval: −2.3%, −0.2%]”, SE from the interval |
| Weiss et al. 2024 [12] | beech forest, NE Germany | 1999–2022 | carabid beetles, abundance | 24-yr pitfall | −0.03149 | 0.01133 | −3.10 | Abstract: “annual rates of −3.1% (95% CI [−5.3, −1])”, SE from the interval |
| Read and cited · not pooled, with the reason | ||||||||
| van Klink et al. 2020 [2, 3] | 41 countries, 1,676 sites | 1925–2018 | terrestrial and freshwater assemblages | 166 studies | −0.01116 | — | −1.11 | A synthesis of overlapping primary data. Pooling it alongside its own inputs would double-count. Used as an external benchmark. |
| Crossley et al. 2020 [6] | 68 US LTER areas | 1975–2019 | arthropods, 5,375 time series | 12 LTER sites | — | — | — | Effect size is change in standard deviations per unit of rescaled time. Not convertible to an annual rate. See [7] for the sampling-history critique. |
| Lister & Garcia 2018 [13] | Luquillo, Puerto Rico | 1976–2013 | arthropod biomass, sticky and sweep | 2 elevations | — | — | — | Reports fold-changes (4–8× sweep, 30–60× sticky), no slope and no interval. Contested by [14] on the same forest. |
| Macgregor et al. 2019 [15] | Great Britain, light traps | 1967–2017 | moth biomass | 50 years | — | — | — | Finds fluctuation without a clear trend; a software error corrupting a quarter of the extraction was corrected in 2021, after which the first, last and peak decades did not differ. |
| Møller 2019 [17] | two transects, Denmark | 1997–2017 | insects on a car windscreen | 1,375 surveys | — | — | — | Retracted in 2026 for multiple data inconsistencies including large-scale duplication. Excluded on those grounds. |
| Haase et al. 2023 [18] | 22 European countries | 1968–2020 | freshwater invertebrate abundance | 1,816 series | +0.01163 | — | +1.17 | Freshwater, so not comparable with terrestrial series; also a synthesis. Used as an external benchmark. |
| Sockman 2025 [19] | one subalpine meadow, Colorado | 2004–2024 | flying insect abundance | 15 seasons | −0.06829 | — | −6.60 | Rate printed without an interval. Quoted in the text as a benchmark for how steep a single well-kept site can be. |
Table 1. Full extraction. Rows in the upper block enter the pool. Rows in the lower block do not, and the last column says why. Every ρ in the upper block is either printed as such in the source, or is a printed log-10 slope multiplied by \(\ln 10\), or is derived from a printed percent-per-year and its printed 95 percent interval. Nothing was read off a figure.
Everything On One Axis
This section is arithmetic. Nothing in it is arguable.
Published trends arrive in at least five currencies: a total percentage change over a stated period, a percentage per year, a percentage per decade, a slope on a base-10 logarithmic scale, and a slope in standard-deviation units per unit of rescaled time. Pooling needs one currency. We use \(\rho\), the natural-log annual rate, defined by
$$N(t) = N_0\, e^{\rho t} \qquad\Longrightarrow\qquad \rho = \frac{1}{t}\ln\!\left(\frac{N(t)}{N_0}\right)$$and converted back for reading as \(100(e^{\rho}-1)\) percent per year or \(100(e^{10\rho}-1)\) percent per decade.
Two conversions were needed. A slope \(b\) reported per year on a base-10 log scale, with standard error \(s\), becomes \(\rho = b\ln 10\) and \(\mathrm{SE}(\rho) = s\ln 10\), where \(\ln 10 = 2.302585\). A rate reported as \(p\) percent per year with a 95 percent interval \([p_{\text{lo}}, p_{\text{hi}}]\) becomes
$$\rho = \ln\!\left(1+\tfrac{p}{100}\right), \qquad \mathrm{SE}(\rho) = \frac{\ln\!\left(1+\tfrac{p_{\text{hi}}}{100}\right) - \ln\!\left(1+\tfrac{p_{\text{lo}}}{100}\right)}{2 \times 1.959964}$$Worked once, for the reader who wants to check a row by hand. Shortall's Hereford biomass slope is −0.01885 per year with a standard error of 0.00410 on the log-10 scale. Multiply both by 2.302585: \(\rho = -0.043404\), \(\mathrm{SE} = 0.009441\). Then \(100(e^{-0.043404}-1) = -4.25\) percent per year. Over the 30 years of that series, \(e^{-0.043404 \times 29} = 0.284\), a 71.6 percent fall.
Two facts about this scale, both of them routinely got wrong in both directions. First, a 6.1 percent annual decline compounds to 76.7 percent over 27 years, so the Krefeld headline and the Krefeld slope are one statement said twice. Second, the gap between −6.11 and −1.11 percent per year looks like nothing on a page and works out at a factor of nine after thirty years, which is the difference between a ruined fauna and a slightly thinner one. Small annual rates are not small.
The three headline claims in §1, converted: −0.0630 for Krefeld, −0.0112 for the 166-survey terrestrial estimate after its erratum, and, for the American null result, a blank where the number should be. That last one resists conversion entirely. Its effect size is change in standard deviations per unit of rescaled time [6], a quantity made dimensionless in the one way that destroys the information anybody would need to get back to a rate per year. Off the axis. We did not put it there.
Pooling, And The Model That Makes It Legal
A fixed-effect meta-analysis assumes every study estimates the same true number and differs from the rest only by sampling error. Weight each study by its precision, \(w_i = 1/\mathrm{SE}_i^2\). Take the weighted mean. Run it on our eleven effects. Out comes −4.26 percent per year, standard error 0.00145, satisfyingly tight.
The tightness is fake. Krefeld reports a standard error of 0.0020 because its model saw 1,503 trap samples drawn from 63 reserves, which is a great deal of information about those 63 reserves and none at all about anywhere else. That one study then carries 52.8 percent of the entire fixed-effect weight, and the pooled answer is mostly the Krefeld answer wearing a meta-analysis costume. Precision within a study says nothing about whether its site resembles anywhere else.
A random-effects model assumes instead that each study estimates its own true value, and that those true values are themselves drawn from a distribution with mean \(\mu\) and between-study variance \(\tau^2\), which is a weaker and far more plausible assumption about the world. The DerSimonian-Laird estimator of that variance [20] is the classical one and the one we implemented:
$$Q = \sum_i w_i (y_i - \hat\mu_{\text{FE}})^2, \qquad C = \sum_i w_i - \frac{\sum_i w_i^2}{\sum_i w_i}, \qquad \hat\tau^2 = \max\!\left(0,\ \frac{Q - (k-1)}{C}\right)$$with revised weights \(w_i^\star = 1/(\mathrm{SE}_i^2 + \hat\tau^2)\), pooled estimate \(\hat\mu = \sum w_i^\star y_i / \sum w_i^\star\), and \(\mathrm{SE}(\hat\mu) = (\sum w_i^\star)^{-1/2}\).
Before we pointed that code at anything real we made it prove itself on data whose answer we already knew. We generated 14 synthetic studies from a known model with true \(\mu = -0.025\) and true \(\tau = 0.015\), pooled them, and got −0.02253 with a 95 percent interval of −0.03054 to −0.01452, and the truth sits inside. Repeating that 2,000 times, the interval covered the true value 91.3 percent of the time and the mean estimate of \(\tau\) came to 0.01425 against a truth of 0.015. Slight under-coverage and slight downward bias in \(\tau\) are documented behaviour for DerSimonian-Laird at small \(k\), so the estimator is working and the code is not failing. We then checked the identity the method guarantees: when \(\hat\tau^2\) comes out at zero, the random-effects weights collapse onto the fixed-effect weights and the two estimates must agree exactly. They agree to 0.000e+00 in the constructed case and in the clamped case. Both checks print at the top of the output file.
ρ = −0.02717, 95% CI −0.04435 to −0.00999. As a rate: −2.68 percent per year (95% CI −4.34 to −0.99), or −23.8 percent per decade (95% CI −35.8 to −9.5). z = −3.10, p = 0.0019.
Between-study variance τ² = 0.000791, so τ = 0.0281. Cochran's Q = 269.96 on 10 degrees of freedom, p = 3.4 × 10−52. I² = 96.3 percent.
95% prediction interval for the next comparable study: −8.14 to +3.10 percent per year.
Figure 2 is the forest plot. Watch the box sizes, which are drawn proportional to the random-effects weight each study actually receives rather than to the precision that study claims for itself. They are all very nearly identical, between 8.36 and 9.66 percent. A large \(\tau^2\) does that: once the model accepts that true effects genuinely differ between sites, knowing one site's effect very precisely stops being much use for estimating the average, and a study with a standard error of 0.0020 gets almost exactly the same say as one with 0.0113.
wRE column; they range only from 8.4 to 9.7 percent, because \(\hat\tau^2\) is large enough that study precision has almost stopped mattering. The diamond is the pooled estimate, −2.68 percent per year. The dashed bar below it is the 95 percent prediction interval, which is the interval a reader should look at if they want to know what a new study might find. It crosses zero comfortably.Ninety-Six Point Three
\(I^2\) is the share of variation sampling error cannot explain [21]. The definition is \((Q - \mathrm{df})/Q\). By convention, anything above 75 percent counts as considerable heterogeneity. Ours is 96.3 percent.
Cochran's \(Q\) is 269.96 on 10 degrees of freedom, which is the sort of number that usually means somebody has pooled apples with lighthouses. If all eleven studies estimated the same true rate, \(Q\) should average 10. The probability of observing 270 is about 3 in 1052. Homogeneity is not marginally rejected here; it is demolished.
So what is the pooled estimate for, if the studies it pools are demonstrably not measuring one common thing?
It summarises the centre of a distribution whose spread is the more interesting quantity, and a centre reported without its spread is a half-written sentence presented as a whole one. \(\hat\tau = 0.0281\) on the log scale puts the standard deviation of true annual rates across monitored sites at about 2.8 percentage points, a number nobody quotes and everybody should. Set that against a mean of 2.68 percent. The spread of the true effects is as large as the effect. Sites are not noisy versions of one common decline; they are doing genuinely different things, and about a sixth of them, on this model, are doing the opposite thing.
The prediction interval makes this concrete. We would give a newspaper that number and no other. Its formula is \(\hat\mu \pm 1.96\sqrt{\hat\tau^2 + \mathrm{SE}(\hat\mu)^2}\), and it answers the question a reader actually has, which concerns the next study rather than the average of the old ones. Ours runs from −8.14 to +3.10 percent per year. A researcher who sets up a new long-term programme somewhere plausible, runs it for twenty years, and finds insects increasing at 2 percent a year has found something entirely compatible with the published literature. So has one who finds a 7 percent annual collapse.
Both of those people will get a press release. Only one of them will get read.
The temptation at this point is moderator analysis: split by continent, by habitat, by taxon, by method, then hunt for the variable that explains the spread. We ran four such splits in §9. None dented \(I^2\) below 87 percent. With eleven effects, moderator hunting generates stories rather than testing them. We are not going to pretend otherwise.
Working Notes, Three Thursdays
Thursday 1 · 12:40 · six present · extraction beginsStarted with Krefeld. Found the slope in four minutes: −0.063, sd 0.002. Then argued for eleven about "sd": posterior standard deviation, or standard error of the coefficient? Settled on the standard error, because a standard deviation of the raw slope estimates across 1,503 samples would be nowhere near that small, and because the paper's own 76.7 percent confidence band of 74.8 to 78.5 percent implies exactly that precision. Checked: \(\ln(0.233)/27 = -0.0539\) using the seasonal mean. Close enough to −0.063, given that the two run over slightly different windows. Kept the printed number. Not our reconstruction.
Thursday 1 · 13:25 · a disagreement, recordedM. wanted the Krefeld study out for being the one everyone has heard of. Fame is not a reason. J. wanted it out because its standard error would dominate any fixed-effect weighting, which is a reason attached to the wrong remedy, the remedy being a random-effects model. Compromise: keep it, report its weight share under both models. Fixed, 52.8 percent. Random, 9.7 percent. Both numbers are in §5.
Thursday 2 · 12:15 · the Shortall problemFound the 2009 Rothamsted suction-trap paper by chasing a citation and immediately had a small crisis, because one of its four traps shows insects going up and the effect is not tiny. Someone asked whether we were allowed to include a positive result. Everyone looked at each other. Including a positive result is the whole point of the exercise, so obviously yes, and that the question got asked out loud at all is the most honest thing in these notes.
Thursday 2 · 13:05 · retraction foundP. was checking the Danish windscreen study [17]. More than 80 percent fewer insects killed on a car over 21 years. Retracted this year. The notice cites multiple data inconsistencies, including duplications affecting large portions of the dataset. We had already copied its numbers into the spreadsheet. Now deleted. Nobody in this room would have caught that by reading the paper; we caught it because a search result carried the word RETRACTION in capitals at the front of the title.
Thursday 3 · 12:50 · the code disagrees with usFirst run of the pooling script gave −4.26 percent per year and we were pleased with it for about five minutes, until someone noticed we had read the fixed-effect line instead of the random-effects line. The random-effects answer is −2.68. Those 1.58 percentage points a year separate a 72 percent loss from a 55 percent loss over thirty years, and the whole gap came from which of two lines of output caught our eye first.
Thursday 3 · 13:30 · closing positionVote on the sentence "insects are declining." Seven for, two abstain, none against. Vote on "insects are declining at about 2.7 percent per year." Two for, three abstain, four against. That gap is the article.
The Funnel Is Not A Lie Detector
Small-study bias is the worry that little studies reach print only when they find something dramatic, while little studies finding nothing sit in a drawer. Where that happens, the low-precision end of the evidence base shifts systematically away from the truth, and every pooled estimate built on it inherits the shift without showing the slightest outward sign of having done so. The standard diagnostic is the funnel plot: effect on one axis, precision on the other. Absent bias, a symmetric inverted funnel. Wide at the bottom where studies are imprecise, narrowing to a point.
Egger's test formalises the eyeballing by regressing each study's standardised effect \(y_i/\mathrm{SE}_i\) on its precision \(1/\mathrm{SE}_i\) and testing whether the intercept differs from zero [22]. Ours does. Intercept +6.87, SE 2.18, \(t = 3.16\) on 9 degrees of freedom, \(p = 0.012\).
Now look at Figure 3 before deciding what that intercept means, because the arithmetic of Egger's test is indifferent to which end of the funnel is doing the tilting.
The asymmetry here carries a sign publication bias does not predict. In the classical pattern the imprecise studies at the bottom of the funnel are the ones shifted towards large effects, because a small study with a small effect never got published. Ours does the reverse. The most extreme effect belongs to the most precise study, sitting alone at the top and well outside the pseudo-confidence region, while the bottom of the funnel holds a positive result at +1.29 percent per year, exactly the finding a publication-bias story says should be missing.
Our reading of that point runs differently. The Krefeld standard error, however correctly computed inside its own model, describes uncertainty about German nature-reserve Malaise catches rather than uncertainty about insect trends anywhere else. Its precision is real, and it is local, and a meta-analysis has no way of telling the two apart from a printed standard error. On a funnel plot that reads as an outlier. In the world it reads as a well-measured place.
We are also obliged to say that Egger's test at \(k = 11\) has very little power, and that its recommended minimum is usually quoted as ten studies, which puts us barely over the line. We ran it because we had said in advance that we would run it, which is the only defensible reason for running a test whose result you already suspect you will disbelieve. We are not leaning on it.
Take One Out, Then Take Out The Duplicates
Leave-one-out is the simplest sensitivity analysis there is: drop each study in turn, re-pool the remaining ten, and see how much the answer moves. If one study is carrying the whole result, an omission that shifts the pool a long way will say so immediately.
Nothing is. Across all eleven omissions the pooled estimate ranges from −3.05 to −2.29 percent per year, and every one of those values sits inside the confidence interval of the full-data estimate. Dropping the Krefeld study, the one carrying half the fixed-effect weight, moves the pooled figure from −2.68 to −2.29 and drops \(I^2\) from 96.3 to 87.0 percent. Substantial, and not decisive. Figure 4 plots all eleven omissions against the full-data interval.
Then the dependence problem, which has two parts. Four of our eleven effects come from one paper, three from another. The pool is not eleven independent pieces of evidence. Worse: the 2025 national butterfly compilation [11] aggregates 35 monitoring programmes across the United States, and the Ohio network analysed separately in 2019 [10] is very likely among them, which makes those two rows partly the same data counted twice. Nothing printed in either paper lets us quantify how much of the American data appears twice. So we flag it. Nine independent datasets at best, not eleven. Collapsing to one effect per paper leaves \(k = 6\) and gives −3.47 percent per year (95% CI −5.65 to −1.41), while collapsing to one effect per locality leaves \(k = 10\) and gives −2.49. Splitting by metric is more interesting. The five biomass series pool to −2.04 percent per year with an interval of −5.53 to +1.42, which crosses zero; the six abundance series pool to −3.14 with an interval of −4.39 to −1.99, which does not. Splitting by continent gives −2.92 for the nine European effects and −1.67 for the two North American ones, a gap that echoes the published finding that continental origin matters a great deal [2] and that rests here on exactly two studies. Treat it as a rhyme rather than a result.
The honest summary of this section: the pooled estimate is stable against every reasonable rearrangement of the same eleven numbers, and that stability tells you nothing at all about whether the eleven numbers are the right eleven.
The Best Argument Against This Article
Suppose the entire literature is an artefact. The case for that is better than it sounds.
Long-term monitoring sites are not chosen at random. Somebody chose them, one site at a time, for reasons that had little to do with sampling theory and a great deal to do with where insects were already known to be. A nature reserve gets designated because it is notably good. A volunteer starts a butterfly transect on the meadow full of butterflies. Funding follows the site where somebody already found a result. A university puts a field station where something is worth studying, then keeps it there for forty years, by which time the station itself is half the reason the place still looks interesting. Selection on abundance is baked into the process by which ecological monitoring comes to exist at all, and it operates on the abundance observed at the moment of selection, which carries a year-specific fluctuation sitting on top of the site's permanent quality.
Regression to the mean does the rest, manufacturing a downward trend out of nothing. If a site was picked partly because year one happened to be a good year, year two should come in worse, purely because the lucky part of year one does not repeat. Fit a straight line through a series that starts on a fluke. You will fit a decline. Nothing out in the world has to change for that line to slope downwards. The critical literature has made this case carefully and repeatedly [18], and the choice of reference year is precisely the point on which the 166-survey synthesis was challenged [4] and defended [5].
So we simulated it, deliberately setting the model up to be unkind to ourselves. Four hundred sites. Twenty-seven years, matching the Krefeld window. Permanent site quality drawn from a normal distribution, standard deviation 1.0 on the log scale. Independent year-to-year noise on top. The true trend is exactly zero. Rank the sites on their first year's observed count, monitor only the top fraction, fit an ordinary least-squares log-linear trend to each surviving series, average the slopes. Figure 5 is the result, at three levels of year-to-year noise.
The mechanism is real and the figure shows it working. At realistic settings it is also far too small to do the job. With year-to-year noise of SD 0.6, roughly a factor of 1.8 between a good year and a bad one, and keeping the top tenth of sites, selection manufactures −0.39 percent per year out of a world with no trend in it. Call that 14.6 percent of our pooled estimate. Push the noise to SD 1.0, keep only the top 2.5 percent, an absurd degree of selection, and you reach −1.36 percent per year, which is the largest artefact this mechanism will produce under any setting anyone has ever proposed out loud. Still only half.
Three further things the simulation shows, one of which cuts against us. Turning the noise down to SD 0.2 nearly abolishes the effect, which confirms regression to the mean as the mechanism rather than some artefact of the simulation we had failed to notice. Adding a real decline of 1 percent per year underneath does not switch selection off; the observed trend becomes −1.39 percent per year at top-tenth selection, so selection inflates true declines as well as inventing false ones. And the effect rises monotonically with selection strength, so anyone arguing that monitoring programmes are sited more selectively than our worst setting holds a coherent position. Just not an evidenced one.
Site selection bias is real. Its mechanism is regression to the mean, and it submits to measurement. Under a model chosen to be unfavourable to us it accounts for roughly 15 percent of the pooled decline, rising to about 50 percent only under settings nobody can defend.
It does not manufacture 76 percent over 27 years. It does inflate whatever real signal exists, which means published rates should be read as upper bounds on the true rate at a randomly chosen place, and it gives one more reason to trust the prediction interval over the point estimate.
The Measurement Problem, Which Is The Actual Story
Set the pooled number aside.
Eleven estimates. One Malaise trap network in German reserves says 6.11 percent per year. One suction trap in Hertfordshire, thirty years long and run by a national institute, says insects went up. These are not competing claims about one quantity, with one of them mistaken. They are accurate reports of what happened in two places, made with two instruments that catch different things, over overlapping but not identical decades, summarised by two statistical models with different ideas about what a trend is.
Consider what the word "insects" is doing in a sentence like "insects are declining at 9 percent per decade." Perhaps five and a half million insect species exist, of which around a million are described. A Malaise trap samples flying insects that fly into mesh, weighted by body mass, so a good year for St Mark's flies swamps everything else in the biomass index [8]. A light trap samples moths and beetles that are phototactic, which is a behaviour, and one that artificial lighting in the surrounding countryside has been altering across the whole period of every one of these time series. A pitfall trap samples ground-active arthropods weighted by how much they walk, so a warm dry summer raises the catch without raising the population by a single beetle. A butterfly transect samples what a volunteer can name on the wing.
Four instruments. Four different functions of the underlying community, each entangled with weather, each entangled with observer behaviour, and not one of them measuring abundance in the sense a reader assumes when they meet the word in a headline. The literature knows this perfectly well; the review that lists seven challenges to interpreting insect declines puts measurement and sampling near the top of the list [16].
The consequence deserves plainer language than it usually gets. When two insect trend studies disagree, start with the instruments. Only afterwards should anyone suppose that one of them is wrong, or that the world itself differs between the two sites, and in this literature the second supposition gets made far too early and far too often. Our \(I^2\) of 96.3 percent gets read as "the sites genuinely differ." Part of it certainly is exactly that, and the Shortall traps are the proof. Some unknown and currently unknowable share of it is instrument, model choice, reference year and the decision about which taxa to count, and no meta-analysis of published summaries can separate those, because the separation would need the raw catches and one common analysis pipeline applied to every one of them.
Which is exactly what the 166-survey synthesis tried to do [2], which is why its authors got the smallest and most defensible number in the whole field, and which is also why they got a technical comment arguing their inclusion criteria and their treatment of the first year were wrong [4], and a response defending both [5]. That exchange is more informative than either paper alone. Two careful groups, the same 166 datasets, a disagreement about when the clock starts. The disagreement was not about insects.
The measurement problem does not block the science. The measurement problem is the science, at its present frontier.
None of which means the declines are imaginary, and we want to be read as saying the opposite. Ten localities, four countries, two metrics and four sampling methods, and the pooled sign is negative, the leave-one-out is stable, and the one positive result in the drawer is not statistically distinguishable from no change. Something is going down, and it is going down in more places than not. What nobody can currently tell you, from this literature, is how fast, over what area, for which insects, or with what confidence, and the confident single figures in circulation are borrowing precision that the underlying measurements do not have.
What We Would Actually Say To A Newspaper
Six sentences, in the order we would say them, with the interval attached to the estimate every time.
- Insect abundance and biomass are declining across most long-term monitoring sites that have been studied, and the direction of that finding is not in serious doubt.
- The best single number we can extract from eleven published estimates is a fall of about 2.7 percent a year, which compounds to roughly a quarter lost per decade, and it has a confidence interval from 1 to 4.3 percent a year that you should quote alongside it or not quote it at all.
- These studies disagree with each other far more than they disagree with zero, \(I^2 = 96.3\) percent, and the interval within which the next study is expected to land runs from an 8.1 percent annual collapse to a 3.1 percent annual increase.
- Freshwater insects in Europe and North America have been going the other way for decades, partly because water got cleaner [2, 18], and any account of insect decline that leaves this out is selling something.
- Site selection is correct in mechanism, and under an unkind model it accounts for around 15 percent of the apparent decline, so published rates are better read as upper bounds than as estimates.
- We are nine students with a spreadsheet and a copy of Python, we could verify eleven effect sizes and no more, and if you are going to print any of this then print the interval and not the point.
Pooled random-effects estimate −2.68 %/yr (95% CI −4.34 to −0.99), k = 11, from 6 papers and 10 monitoring datasets, of which two share data. I² = 96.3%, τ = 0.028, prediction interval −8.14 to +3.10 %/yr.
Confidence in the sign: high. Confidence in the magnitude: low, and deliberately so. A school club extracted this from printed abstracts and tables; it is not a systematic review, and the pooled point estimate should be read as one more entry in the drawer rather than as a summary of the field.
The calculation lives in insect-decline-meta.py, which uses nothing outside the Python standard library so that anyone can run it, and which prints its own validation checks before it prints a single result. The interactive model lets you throw studies out of the pool by clicking them, watch \(I^2\) move, and turn the selection-bias dial yourself, which is a faster route into this article than reading it twice. Think our pooled number too high? Remove the studies you object to. Then watch the heterogeneity.
It stays above 85 percent no matter which studies you throw out, which is a more durable result than anything else in this article. That is the finding.
References
- Hallmann, C. A., Sorg, M., Jongejans, E., Siepel, H., Hofland, N., Schwan, H., Stenmans, W., Müller, A., Sumser, H., Hörren, T., Goulson, D. & de Kroon, H. (2017). More than 75 percent decline over 27 years in total flying insect biomass in protected areas. PLoS ONE 12(10), e0185809. doi:10.1371/journal.pone.0185809
- van Klink, R., Bowler, D. E., Gongalsky, K. B., Swengel, A. B., Gentile, A. & Chase, J. M. (2020). Meta-analysis reveals declines in terrestrial but increases in freshwater insect abundances. Science 368, 417–420. doi:10.1126/science.aax9931
- van Klink, R., Bowler, D. E., Gongalsky, K. B., Swengel, A. B., Gentile, A. & Chase, J. M. (2020). Erratum for the Report “Meta-analysis reveals declines in terrestrial but increases in freshwater insect abundances”. Science 370, eabf1915. doi:10.1126/science.abf1915 [corrects the terrestrial estimate to −1.11% per year]
- Desquilbet, M., Gaume, L., Grippa, M., Céréghino, R., Humbert, J.-F., Bonmatin, J.-M., Cornillon, P.-A., Maes, D., Van Dyck, H. & Goulson, D. (2020). Comment on “Meta-analysis reveals declines in terrestrial but increases in freshwater insect abundances”. Science 370, eabd8947. doi:10.1126/science.abd8947
- van Klink, R., Bowler, D. E., Gongalsky, K. B., Swengel, A. B., Gentile, A. & Chase, J. M. (2020). Response to Comment on “Meta-analysis reveals declines in terrestrial but increases in freshwater insect abundances”. Science 370, eabe0760. doi:10.1126/science.abe0760
- Crossley, M. S., Meier, A. R., Baldwin, E. M., Berry, L. L., Crenshaw, L. C., Hartman, G. L., Lagos-Kutz, D., Nichols, D. H., Patel, K., Varriano, S., Snyder, W. E. & Moran, M. D. (2020). No net insect abundance and diversity declines across US Long Term Ecological Research sites. Nature Ecology & Evolution 4, 1368–1376. doi:10.1038/s41559-020-1269-4
- Welti, E. A. R., Joern, A., Ellison, A. M., Lightfoot, D. C., Record, S., Rodenhouse, N., Stanley, E. H. & Kaspari, M. (2021). Studies of insect temporal trends must account for the complex sampling histories inherent to many long-term monitoring efforts. Nature Ecology & Evolution 5, 589–591. doi:10.1038/s41559-021-01424-0
- Shortall, C. R., Moore, A., Smith, E., Hall, M. J., Woiwod, I. P. & Harrington, R. (2009). Long-term changes in the abundance of flying insects. Insect Conservation and Diversity 2, 251–260. doi:10.1111/j.1752-4598.2009.00062.x
- Hallmann, C. A., Zeegers, T., van Klink, R., Vermeulen, R., van Wielink, P., Spijkers, H., van Deijk, J., van Steenis, W. & Jongejans, E. (2020). Declining abundance of beetles, moths and caddisflies in the Netherlands. Insect Conservation and Diversity 13, 127–139. doi:10.1111/icad.12377
- Wepprich, T., Adrion, J. R., Ries, L., Wiedmann, J. & Haddad, N. M. (2019). Butterfly abundance declines over 20 years of systematic monitoring in Ohio, USA. PLoS ONE 14(7), e0216270. doi:10.1371/journal.pone.0216270
- Edwards, C. B., Zipkin, E. F., Henry, E. H., Haddad, N. M., Forister, M. L. et al. (2025). Rapid butterfly declines across the United States during the 21st century. Science 387, 1090–1094. doi:10.1126/science.adp4671
- Weiss, F., von Wehrden, H. & Linde, A. (2024). Long-term drought triggers severe declines in carabid beetles in a temperate forest. Ecography 2024, e07020. doi:10.1111/ecog.07020
- Lister, B. C. & Garcia, A. (2018). Climate-driven declines in arthropod abundance restructure a rainforest food web. Proceedings of the National Academy of Sciences 115, E10397–E10406. doi:10.1073/pnas.1722477115
- Schowalter, T. D., Pandey, M., Presley, S. J., Willig, M. R. & Zimmerman, J. K. (2021). Arthropods are not declining but are responsive to disturbance in the Luquillo Experimental Forest, Puerto Rico. Proceedings of the National Academy of Sciences 118, e2002556117. doi:10.1073/pnas.2002556117
- Macgregor, C. J., Williams, J. H., Bell, J. R. & Thomas, C. D. (2019). Moth biomass has fluctuated over 50 years in Britain but lacks a clear trend. Nature Ecology & Evolution 3, 1645–1649. doi:10.1038/s41559-019-1028-6 [see Author Correction, Nature Ecology & Evolution 5, 865–883 (2021), doi:10.1038/s41559-021-01449-5]
- Didham, R. K., Basset, Y., Collins, C. M., Leather, S. R., Littlewood, N. A., Menz, M. H. M., Müller, J., Packer, L., Saunders, M. E., Schönrogge, K., Stewart, A. J. A., Yanoviak, S. P. & Hassall, C. (2020). Interpreting insect declines: seven challenges and a way forward. Insect Conservation and Diversity 13, 103–114. doi:10.1111/icad.12408
- Møller, A. P. (2019). Parallel declines in abundance of insects and insectivorous birds in Denmark over 22 years. Ecology and Evolution 9, 6581–6587. doi:10.1002/ece3.5236 [RETRACTED 2026; retraction notice doi:10.1002/ece3.73864]
- Haase, P., Bowler, D. E., Baker, N. J., Bonada, N., Domisch, S. et al. (2023). The recovery of European freshwater biodiversity has come to a halt. Nature 620, 582–588. doi:10.1038/s41586-023-06400-1
- Sockman, K. W. (2025). Long-term decline in montane insects under warming summers. Ecology 106, e70187. doi:10.1002/ecy.70187
- DerSimonian, R. & Laird, N. (1986). Meta-analysis in clinical trials. Controlled Clinical Trials 7, 177–188. doi:10.1016/0197-2456(86)90046-2
- Higgins, J. P. T. & Thompson, S. G. (2002). Quantifying heterogeneity in a meta-analysis. Statistics in Medicine 21, 1539–1558. doi:10.1002/sim.1186
- Egger, M., Davey Smith, G., Schneider, M. & Minder, C. (1997). Bias in meta-analysis detected by a simple, graphical test. BMJ 315, 629–634. doi:10.1136/bmj.315.7109.629