Science Journaling Club Founded 2024

REVIEW · META-ANALYSIS · ECOLOGY

The Insect Decline Literature Disagrees With Itself, and We Wanted to Know How Badly

Written jointly by the Science Journaling Club

Review · Peer-edited by the club review board · LaTeX source · Our calculation · Interactive model

Abstract One German study reported a 76.7 percent fall in flying-insect biomass over 27 years [1]. A synthesis of 166 long-term surveys put terrestrial declines near 9 percent per decade while freshwater insects rose about 11 percent per decade, and an erratum later steepened the terrestrial figure [2, 3]. More than 5,300 American time series gave net trends indistinguishable from zero [6], which is a statement about time series and about the units they were expressed in, though almost nobody repeated it that way. All three are competent work. We extracted 11 effect sizes from 6 papers covering 10 monitoring datasets, traced every one to a printed slope or a printed interval in its source, converted them to a common log-linear annual rate, and pooled them under a DerSimonian-Laird random-effects model. The pool comes out at −2.68 percent per year (95% CI −4.34 to −0.99), or −23.8 percent per decade, and the spread around it is the part worth printing: I² = 96.3 percent, with Cochran's Q at 269.96 on 10 degrees of freedom. That second pair of numbers matters more than the first. A 95% prediction interval for the next comparable study runs from −8.14 to +3.10 percent per year, so a new paper reporting insects on the rise would sit comfortably inside everything published to date. We also simulate the objection that monitoring programmes began at unusually good sites, where regression to the mean manufactures a decline nobody caused; under a deliberately unkind version of that model the artefact accounts for roughly 15 percent of our pooled estimate. Real and quantified, and nowhere near large enough to carry the objection. We are a school club reading printed tables. Weight the pooled number below what its confidence interval implies. Weight the heterogeneity above it.

Two Numbers That Cannot Both Be The Headline

The problem, in one paragraph.

In October 2017 a team working with amateur entomologists in Krefeld published Malaise-trap biomass from 63 German nature reserves and reported a 76.7 percent fall over 27 years, with a mid-summer figure of 81.6 percent [1]. A Malaise trap is a tent of fine mesh that flying insects blunder into and then crawl upwards into a collecting jar, and it does not care in the least what it catches. That indifference is the instrument's one virtue. The number went around the world in a day. In April 2020 a different team compiled 166 long-term surveys across 1,676 sites and reported a terrestrial decline of about 0.92 percent per year, roughly 9 percent per decade, with freshwater insects moving the other way at about +11 percent per decade [2]. Four months later, in August 2020, a third team pulled more than 5,300 arthropod time series out of the United States Long Term Ecological Research network and found net trends that were, in their words, generally indistinguishable from zero [6].

None of those three papers is bad. We read all of them, twice in two cases, and found the reasoning in each one about as careful as the reasoning in the other two. We would sign off on all of them.

And yet the first says one thing, the second says a smaller version of that thing in one place and the opposite of it in another, and the third, drawing on an archive that had kept the raw counts for decades, says nothing much is happening at all. The public conversation settled the matter the way it settles most matters of this shape, by picking whichever number suited the argument already under way. Press releases got the 76 percent. Columnists who dislike environmental alarm got the zero. Almost nobody quoted the middle estimate, the one built from the most data, because a 9 percent decline per decade makes a footnote rather than a headline.

We got tired of this. So we did the arithmetic ourselves.

What follows is a meta-analysis, which is to say a study whose data are other studies, carrying every error those studies carried plus a few of its own. We put published trend estimates onto a single scale, pooled them under a model that lets the true effect differ between sites, and then spent most of our effort on the diagnostics that say whether the pooling meant anything at all. Figure 1 shows why the disagreement is not cosmetic. Five published annual rates, run forward for thirty years, land anywhere between an 85 percent loss and a 38 percent gain, which is not a disagreement about magnitude but a disagreement about direction.

0 20 40 60 80 100 120 percent of starting count years of monitoring 0 10 20 30 same literature five published rates Krefeld Malaise biomass [1] · −6.11%/yr same pool, fixed-effect model · −4.26%/yr our random-effects pool · −2.68%/yr 166-survey terrestrial [3] · −1.11%/yr freshwater [2] · +1.08%/yr
Figure 1. Five annual rates of change, each taken from a published estimate and run forward for thirty years under \(N(t) = N_0 e^{\rho t}\). Solid deep blue: the Krefeld Malaise biomass rate of −6.11 percent per year [1]. Solid mid blue: our own random-effects pool, −2.68 percent per year. Dashed mid blue: what a fixed-effect model would have given, −4.26 percent per year, shown because it is wrong and it is instructive to see how wrong. Solid pale blue: the 166-survey terrestrial estimate after its erratum, −1.11 percent per year [3]. Dashed orange: the freshwater estimate from the same paper, +1.08 percent per year [2]. Thirty years later the top curve and the bottom curve differ by a factor of nine.

What A School Club Can And Cannot Do To A Literature

Before anyone quotes this, we should be exact about what it is, because a school club's arithmetic and a systematic review are different objects, and only one of them is on offer here.

A systematic review has a pre-registered search string, two people screening every title independently, a documented reason for every exclusion, and usually a request to the original authors for data that was never printed. We have none of that. Our search procedure was to chase reference lists until they repeated. Our screening procedure was six people at a table arguing, which is a worse procedure than two independent screeners and a better one than nobody noticing. Our data extraction was one person reading a table aloud while a second person typed, then the two swapping roles and doing the whole thing again to see whether the numbers came out the same.

That gives us a convenience sample. That sample is this article's worst flaw. No amount of correct arithmetic downstream repairs it. If the papers that came to mind first are the papers that made news, and papers make news by finding large effects, then the pool is tilted before we start. We test for that in §8. We do not entirely like the answer.

Against that, one rule we did keep, and kept strictly. Every effect size in the pool traces to a printed number in its source. Not read off a figure with a ruler, and not imputed from a percentage plus a guessed duration, which is the shortcut that would have let three more papers into the drawer. A study got into the pool if and only if it printed a log-linear annual trend coefficient with a standard error, or printed a percent-per-year with a 95 percent interval from which a standard error could be recovered by arithmetic that any reader can repeat. The provenance column in Table 1 quotes the sentence that carries the number, or names the table and the cell the number sits in, so a reader can find it without asking us anything. Check every row against its source in an afternoon. Tell us if we misread one.

We fixed the second rule before looking at any numbers, which matters more than it sounds, because a rule written afterwards will quietly select the effects that make the story tidier and nobody in the room will notice themselves doing it. One effect per independent monitoring dataset, always the paper's own headline series. We took whichever order the paper itself led with. Never the largest decline on offer. Two papers still contribute more than one effect each, because they hold genuinely separate monitoring programmes at separate sites, and §9 re-runs the whole pool with those collapsed to one apiece.

Third, and worst: we are not measuring insects. We are measuring published measurements of insects, a different quantity with an error structure of its own, and that difference widens until it becomes the whole subject by §11.

The Drawer

Eleven specimens. Ten localities, six papers, four countries, two metrics. Forty-nine years of calendar time separate the earliest start from the latest finish. Pinned and labelled below. A meta-analysis dataset is a collection of things other people caught, mounted by other people again, and a drawer is the honest way to display one.

Drawer 1 · pooled seriesdetermination: log-linear annual rate

Hallmann et al. 2017

Locality
63 protected areas, lowland Germany
Duration
1989–2016, 27 years, 1,503 trap samples
Taxa
all flying insects, Malaise-trap dry biomass
Trend
ρ = −0.0630 ± 0.0020 · −6.11 %/yr

Shortall et al. 2009 · Hereford

Locality
Hereford, western England
Duration
1973–2002, 30 annual indices
Taxa
aerial insects, 12.2 m suction trap, biomass
Trend
ρ = −0.0434 ± 0.0094 · −4.25 %/yr

Shortall et al. 2009 · Rothamsted

Locality
Harpenden, Hertfordshire
Duration
1973–2002, 30 annual indices
Taxa
aerial insects, 12.2 m suction trap, biomass
Trend
ρ = +0.0128 ± 0.0094 · +1.29 %/yr

Shortall et al. 2009 · Starcross

Locality
Starcross, Devon
Duration
1973–2002, 30 annual indices
Taxa
aerial insects, 12.2 m suction trap, biomass
Trend
ρ = −0.0050 ± 0.0061 · −0.50 %/yr

Shortall et al. 2009 · Wye

Locality
Wye, Kent
Duration
1973–2002, 30 annual indices
Taxa
aerial insects, 12.2 m suction trap, biomass
Trend
ρ = −0.0030 ± 0.0061 · −0.30 %/yr

Hallmann et al. 2020 · Kaaistoep

Locality
De Kaaistoep, North Brabant, Netherlands
Duration
1997–2017, 497 trapping nights
Taxa
macro-moths at light, 54,492 individuals
Trend
ρ = −0.0400 ± 0.0060 · −3.92 %/yr

Hallmann et al. 2020 · Kaaistoep

Locality
De Kaaistoep, North Brabant, Netherlands
Duration
1997–2017, 572 trapping nights
Taxa
beetles at light, 257,793 individuals
Trend
ρ = −0.0480 ± 0.0100 · −4.69 %/yr

Hallmann et al. 2020 · Wijster

Locality
Wijster, Drenthe, Netherlands
Duration
1986–2016, 26 sampled years, 48 pitfalls
Taxa
ground beetles, 264,986 individuals, 156 species
Trend
ρ = −0.0440 ± 0.0060 · −4.30 %/yr

Wepprich et al. 2019

Locality
Ohio, USA, 104 monitored sites
Duration
1996–2016, 24,405 volunteer surveys
Taxa
butterflies, 81 species with estimable trends
Trend
ρ = −0.0200 ± 0.0050 · −1.98 %/yr

Edwards et al. 2025

Locality
contiguous USA, 2,478 locations, 35 programmes
Duration
2000–2020, 76,957 surveys
Taxa
butterflies, 554 species, 12.6 M individuals
Trend
ρ = −0.0131 ± 0.0054 · −1.30 %/yr

Weiss et al. 2024

Locality
old beech forest, north-east Germany
Duration
1999–2022, 24-year pitfall series
Taxa
carabid beetles, total abundance
Trend
ρ = −0.0315 ± 0.0113 · −3.10 %/yr

Look at the four Shortall labels [8]. Same country, same instrument, same protocol, same three decades, same team doing the counting, and four answers that a reader would take for four different continents. Hereford loses biomass at 4.25 percent a year. Rothamsted gains it at 1.29 percent a year. Starcross and Wye do essentially nothing. Anyone wanting a single demonstration that the insect trend literature measures something spatially wild rather than a planetary constant can stop reading here, and might note that the demonstration was printed eight years before the Krefeld paper that started the argument.

Table 1 is the extraction sheet, and the only part of this article that would survive our conclusions being wrong. Everything above depends on it, so it gets printed here in full rather than deposited in a repository nobody opens, and the lower block lists the work we read, cited and deliberately did not pool, with the reason alongside each row.

Study Locality Span Taxa / metric Sample ρ yr−1 SE %/yr Where the number is printed
Pooled · slope with standard error, or rate with a 95% interval
Hallmann et al. 2017 [1]63 reserves, Germany1989–2016 flying insects, Malaise biomass1,503 samples −0.063000.00200−6.11 Results: “annual trend coefficient = −0.063, sd = 0.002, i.e. 6.1% annual decline”
Shortall et al. 2009 [8]Hereford, England1973–2002 aerial biomass, suction trap30 indices −0.043400.00944−4.25 Table 1, row “Biomass Hereford”: slope −0.01885, SE 0.00410 (log10), × ln 10
Shortall et al. 2009 [8]Harpenden, England1973–2002 aerial biomass, suction trap30 indices +0.012830.00942+1.29 Table 1, row “Biomass Rothamsted”: slope +0.00557, SE 0.00409 (log10)
Shortall et al. 2009 [8]Starcross, England1973–2002 aerial biomass, suction trap30 indices −0.005000.00610−0.50 Table 1, row “Biomass Starcross”: slope −0.00217, SE 0.00265 (log10)
Shortall et al. 2009 [8]Wye, England1973–2002 aerial biomass, suction trap30 indices −0.002970.00610−0.30 Table 1, row “Biomass Wye”: slope −0.00129, SE 0.00265 (log10)
Hallmann et al. 2020 [9]De Kaaistoep, NL1997–2017 macro-moths at light497 nights −0.040000.00600−3.92 Table 2, Lepidoptera: ρ −0.040 (0.006)
Hallmann et al. 2020 [9]De Kaaistoep, NL1997–2017 beetles at light572 nights −0.048000.01000−4.69 Table 2, Coleoptera: ρ −0.048 (0.010)
Hallmann et al. 2020 [9]Wijster, NL1986–2016 ground beetles, 48 pitfalls264,986 ind. −0.044000.00600−4.30 Results: “ρ = −0.044, se = 0.006, P < 0.001, 4% decline per year”
Wepprich et al. 2019 [10]Ohio, USA1996–2016 butterflies, 81 species24,405 surveys −0.020000.00500−1.98 Results: “declined at an annual rate of 2.0% (β1 = −0.020, std. err. 0.005, p < 0.001)”
Edwards et al. 2025 [11]contiguous USA2000–2020 butterflies, 554 species76,957 surveys −0.013090.00543−1.30 “a rate of 1.3% annually [95% confidence interval: −2.3%, −0.2%]”, SE from the interval
Weiss et al. 2024 [12]beech forest, NE Germany1999–2022 carabid beetles, abundance24-yr pitfall −0.031490.01133−3.10 Abstract: “annual rates of −3.1% (95% CI [−5.3, −1])”, SE from the interval
Read and cited · not pooled, with the reason
van Klink et al. 2020 [2, 3]41 countries, 1,676 sites1925–2018 terrestrial and freshwater assemblages166 studies −0.01116−1.11 A synthesis of overlapping primary data. Pooling it alongside its own inputs would double-count. Used as an external benchmark.
Crossley et al. 2020 [6]68 US LTER areas1975–2019 arthropods, 5,375 time series12 LTER sites Effect size is change in standard deviations per unit of rescaled time. Not convertible to an annual rate. See [7] for the sampling-history critique.
Lister & Garcia 2018 [13]Luquillo, Puerto Rico1976–2013 arthropod biomass, sticky and sweep2 elevations Reports fold-changes (4–8× sweep, 30–60× sticky), no slope and no interval. Contested by [14] on the same forest.
Macgregor et al. 2019 [15]Great Britain, light traps1967–2017 moth biomass50 years Finds fluctuation without a clear trend; a software error corrupting a quarter of the extraction was corrected in 2021, after which the first, last and peak decades did not differ.
Møller 2019 [17]two transects, Denmark1997–2017 insects on a car windscreen1,375 surveys Retracted in 2026 for multiple data inconsistencies including large-scale duplication. Excluded on those grounds.
Haase et al. 2023 [18]22 European countries1968–2020 freshwater invertebrate abundance1,816 series +0.01163+1.17 Freshwater, so not comparable with terrestrial series; also a synthesis. Used as an external benchmark.
Sockman 2025 [19]one subalpine meadow, Colorado2004–2024 flying insect abundance15 seasons −0.06829−6.60 Rate printed without an interval. Quoted in the text as a benchmark for how steep a single well-kept site can be.

Table 1. Full extraction. Rows in the upper block enter the pool. Rows in the lower block do not, and the last column says why. Every ρ in the upper block is either printed as such in the source, or is a printed log-10 slope multiplied by \(\ln 10\), or is derived from a printed percent-per-year and its printed 95 percent interval. Nothing was read off a figure.

Everything On One Axis

This section is arithmetic. Nothing in it is arguable.

Published trends arrive in at least five currencies: a total percentage change over a stated period, a percentage per year, a percentage per decade, a slope on a base-10 logarithmic scale, and a slope in standard-deviation units per unit of rescaled time. Pooling needs one currency. We use \(\rho\), the natural-log annual rate, defined by

$$N(t) = N_0\, e^{\rho t} \qquad\Longrightarrow\qquad \rho = \frac{1}{t}\ln\!\left(\frac{N(t)}{N_0}\right)$$

and converted back for reading as \(100(e^{\rho}-1)\) percent per year or \(100(e^{10\rho}-1)\) percent per decade.

Two conversions were needed. A slope \(b\) reported per year on a base-10 log scale, with standard error \(s\), becomes \(\rho = b\ln 10\) and \(\mathrm{SE}(\rho) = s\ln 10\), where \(\ln 10 = 2.302585\). A rate reported as \(p\) percent per year with a 95 percent interval \([p_{\text{lo}}, p_{\text{hi}}]\) becomes

$$\rho = \ln\!\left(1+\tfrac{p}{100}\right), \qquad \mathrm{SE}(\rho) = \frac{\ln\!\left(1+\tfrac{p_{\text{hi}}}{100}\right) - \ln\!\left(1+\tfrac{p_{\text{lo}}}{100}\right)}{2 \times 1.959964}$$

Worked once, for the reader who wants to check a row by hand. Shortall's Hereford biomass slope is −0.01885 per year with a standard error of 0.00410 on the log-10 scale. Multiply both by 2.302585: \(\rho = -0.043404\), \(\mathrm{SE} = 0.009441\). Then \(100(e^{-0.043404}-1) = -4.25\) percent per year. Over the 30 years of that series, \(e^{-0.043404 \times 29} = 0.284\), a 71.6 percent fall.

Two facts about this scale, both of them routinely got wrong in both directions. First, a 6.1 percent annual decline compounds to 76.7 percent over 27 years, so the Krefeld headline and the Krefeld slope are one statement said twice. Second, the gap between −6.11 and −1.11 percent per year looks like nothing on a page and works out at a factor of nine after thirty years, which is the difference between a ruined fauna and a slightly thinner one. Small annual rates are not small.

The three headline claims in §1, converted: −0.0630 for Krefeld, −0.0112 for the 166-survey terrestrial estimate after its erratum, and, for the American null result, a blank where the number should be. That last one resists conversion entirely. Its effect size is change in standard deviations per unit of rescaled time [6], a quantity made dimensionless in the one way that destroys the information anybody would need to get back to a rate per year. Off the axis. We did not put it there.

Pooling, And The Model That Makes It Legal

A fixed-effect meta-analysis assumes every study estimates the same true number and differs from the rest only by sampling error. Weight each study by its precision, \(w_i = 1/\mathrm{SE}_i^2\). Take the weighted mean. Run it on our eleven effects. Out comes −4.26 percent per year, standard error 0.00145, satisfyingly tight.

The tightness is fake. Krefeld reports a standard error of 0.0020 because its model saw 1,503 trap samples drawn from 63 reserves, which is a great deal of information about those 63 reserves and none at all about anywhere else. That one study then carries 52.8 percent of the entire fixed-effect weight, and the pooled answer is mostly the Krefeld answer wearing a meta-analysis costume. Precision within a study says nothing about whether its site resembles anywhere else.

A random-effects model assumes instead that each study estimates its own true value, and that those true values are themselves drawn from a distribution with mean \(\mu\) and between-study variance \(\tau^2\), which is a weaker and far more plausible assumption about the world. The DerSimonian-Laird estimator of that variance [20] is the classical one and the one we implemented:

$$Q = \sum_i w_i (y_i - \hat\mu_{\text{FE}})^2, \qquad C = \sum_i w_i - \frac{\sum_i w_i^2}{\sum_i w_i}, \qquad \hat\tau^2 = \max\!\left(0,\ \frac{Q - (k-1)}{C}\right)$$

with revised weights \(w_i^\star = 1/(\mathrm{SE}_i^2 + \hat\tau^2)\), pooled estimate \(\hat\mu = \sum w_i^\star y_i / \sum w_i^\star\), and \(\mathrm{SE}(\hat\mu) = (\sum w_i^\star)^{-1/2}\).

Before we pointed that code at anything real we made it prove itself on data whose answer we already knew. We generated 14 synthetic studies from a known model with true \(\mu = -0.025\) and true \(\tau = 0.015\), pooled them, and got −0.02253 with a 95 percent interval of −0.03054 to −0.01452, and the truth sits inside. Repeating that 2,000 times, the interval covered the true value 91.3 percent of the time and the mean estimate of \(\tau\) came to 0.01425 against a truth of 0.015. Slight under-coverage and slight downward bias in \(\tau\) are documented behaviour for DerSimonian-Laird at small \(k\), so the estimator is working and the code is not failing. We then checked the identity the method guarantees: when \(\hat\tau^2\) comes out at zero, the random-effects weights collapse onto the fixed-effect weights and the two estimates must agree exactly. They agree to 0.000e+00 in the constructed case and in the clamped case. Both checks print at the top of the output file.

Determination · the pooled estimate

ρ = −0.02717, 95% CI −0.04435 to −0.00999. As a rate: −2.68 percent per year (95% CI −4.34 to −0.99), or −23.8 percent per decade (95% CI −35.8 to −9.5). z = −3.10, p = 0.0019.

Between-study variance τ² = 0.000791, so τ = 0.0281. Cochran's Q = 269.96 on 10 degrees of freedom, p = 3.4 × 10−52. I² = 96.3 percent.

95% prediction interval for the next comparable study: −8.14 to +3.10 percent per year.

Figure 2 is the forest plot. Watch the box sizes, which are drawn proportional to the random-effects weight each study actually receives rather than to the precision that study claims for itself. They are all very nearly identical, between 8.36 and 9.66 percent. A large \(\tau^2\) does that: once the model accepts that true effects genuinely differ between sites, knowing one site's effect very precisely stops being much use for estimating the average, and a study with a standard error of 0.0020 gets almost exactly the same say as one with 0.0113.

STUDY / DATASET wRE %/yr [95% CI] Hallmann 2017 · Germany 9.7 −6.11 [−6.47, −5.74] Shortall 2009 · Hereford 8.7 −4.25 [−6.00, −2.46] Shortall 2009 · Rothamsted 8.7 +1.29 [−0.56, +3.18] Shortall 2009 · Starcross 9.3 −0.50 [−1.68, +0.70] Shortall 2009 · Wye 9.3 −0.30 [−1.48, +0.90] Hallmann 2020 · moths 9.3 −3.92 [−5.04, −2.78] Hallmann 2020 · beetles 8.6 −4.69 [−6.54, −2.80] Hallmann 2020 · carabids 9.3 −4.30 [−5.42, −3.17] Wepprich 2019 · Ohio 9.4 −1.98 [−2.94, −1.01] Edwards 2025 · US 9.4 −1.30 [−2.34, −0.24] Weiss 2024 · NE Germany 8.4 −3.10 [−5.23, −0.92] POOLED (random effects, k = 11) −2.68 [−4.34, −0.99] 95% prediction interval −8.14 to +3.10 −9 −7 −5 −3 −1 0 +2 +4 change in total abundance or biomass, percent per year
Figure 2. Forest plot of the eleven pooled effects, with 95 percent confidence intervals, ordered as in Table 1. Box areas are proportional to the random-effects weight, printed in the wRE column; they range only from 8.4 to 9.7 percent, because \(\hat\tau^2\) is large enough that study precision has almost stopped mattering. The diamond is the pooled estimate, −2.68 percent per year. The dashed bar below it is the 95 percent prediction interval, which is the interval a reader should look at if they want to know what a new study might find. It crosses zero comfortably.

Ninety-Six Point Three

\(I^2\) is the share of variation sampling error cannot explain [21]. The definition is \((Q - \mathrm{df})/Q\). By convention, anything above 75 percent counts as considerable heterogeneity. Ours is 96.3 percent.

Cochran's \(Q\) is 269.96 on 10 degrees of freedom, which is the sort of number that usually means somebody has pooled apples with lighthouses. If all eleven studies estimated the same true rate, \(Q\) should average 10. The probability of observing 270 is about 3 in 1052. Homogeneity is not marginally rejected here; it is demolished.

So what is the pooled estimate for, if the studies it pools are demonstrably not measuring one common thing?

It summarises the centre of a distribution whose spread is the more interesting quantity, and a centre reported without its spread is a half-written sentence presented as a whole one. \(\hat\tau = 0.0281\) on the log scale puts the standard deviation of true annual rates across monitored sites at about 2.8 percentage points, a number nobody quotes and everybody should. Set that against a mean of 2.68 percent. The spread of the true effects is as large as the effect. Sites are not noisy versions of one common decline; they are doing genuinely different things, and about a sixth of them, on this model, are doing the opposite thing.

The prediction interval makes this concrete. We would give a newspaper that number and no other. Its formula is \(\hat\mu \pm 1.96\sqrt{\hat\tau^2 + \mathrm{SE}(\hat\mu)^2}\), and it answers the question a reader actually has, which concerns the next study rather than the average of the old ones. Ours runs from −8.14 to +3.10 percent per year. A researcher who sets up a new long-term programme somewhere plausible, runs it for twenty years, and finds insects increasing at 2 percent a year has found something entirely compatible with the published literature. So has one who finds a 7 percent annual collapse.

Both of those people will get a press release. Only one of them will get read.

The temptation at this point is moderator analysis: split by continent, by habitat, by taxon, by method, then hunt for the variable that explains the spread. We ran four such splits in §9. None dented \(I^2\) below 87 percent. With eleven effects, moderator hunting generates stories rather than testing them. We are not going to pretend otherwise.

Working Notes, Three Thursdays

Thursday 1 · 12:40 · six present · extraction begins

Started with Krefeld. Found the slope in four minutes: −0.063, sd 0.002. Then argued for eleven about "sd": posterior standard deviation, or standard error of the coefficient? Settled on the standard error, because a standard deviation of the raw slope estimates across 1,503 samples would be nowhere near that small, and because the paper's own 76.7 percent confidence band of 74.8 to 78.5 percent implies exactly that precision. Checked: \(\ln(0.233)/27 = -0.0539\) using the seasonal mean. Close enough to −0.063, given that the two run over slightly different windows. Kept the printed number. Not our reconstruction.

Thursday 1 · 13:25 · a disagreement, recorded

M. wanted the Krefeld study out for being the one everyone has heard of. Fame is not a reason. J. wanted it out because its standard error would dominate any fixed-effect weighting, which is a reason attached to the wrong remedy, the remedy being a random-effects model. Compromise: keep it, report its weight share under both models. Fixed, 52.8 percent. Random, 9.7 percent. Both numbers are in §5.

Thursday 2 · 12:15 · the Shortall problem

Found the 2009 Rothamsted suction-trap paper by chasing a citation and immediately had a small crisis, because one of its four traps shows insects going up and the effect is not tiny. Someone asked whether we were allowed to include a positive result. Everyone looked at each other. Including a positive result is the whole point of the exercise, so obviously yes, and that the question got asked out loud at all is the most honest thing in these notes.

Thursday 2 · 13:05 · retraction found

P. was checking the Danish windscreen study [17]. More than 80 percent fewer insects killed on a car over 21 years. Retracted this year. The notice cites multiple data inconsistencies, including duplications affecting large portions of the dataset. We had already copied its numbers into the spreadsheet. Now deleted. Nobody in this room would have caught that by reading the paper; we caught it because a search result carried the word RETRACTION in capitals at the front of the title.

Thursday 3 · 12:50 · the code disagrees with us

First run of the pooling script gave −4.26 percent per year and we were pleased with it for about five minutes, until someone noticed we had read the fixed-effect line instead of the random-effects line. The random-effects answer is −2.68. Those 1.58 percentage points a year separate a 72 percent loss from a 55 percent loss over thirty years, and the whole gap came from which of two lines of output caught our eye first.

Thursday 3 · 13:30 · closing position

Vote on the sentence "insects are declining." Seven for, two abstain, none against. Vote on "insects are declining at about 2.7 percent per year." Two for, three abstain, four against. That gap is the article.

The Funnel Is Not A Lie Detector

Small-study bias is the worry that little studies reach print only when they find something dramatic, while little studies finding nothing sit in a drawer. Where that happens, the low-precision end of the evidence base shifts systematically away from the truth, and every pooled estimate built on it inherits the shift without showing the slightest outward sign of having done so. The standard diagnostic is the funnel plot: effect on one axis, precision on the other. Absent bias, a symmetric inverted funnel. Wide at the bottom where studies are imprecise, narrowing to a point.

Egger's test formalises the eyeballing by regressing each study's standardised effect \(y_i/\mathrm{SE}_i\) on its precision \(1/\mathrm{SE}_i\) and testing whether the intercept differs from zero [22]. Ours does. Intercept +6.87, SE 2.18, \(t = 3.16\) on 9 degrees of freedom, \(p = 0.012\).

Now look at Figure 3 before deciding what that intercept means, because the arithmetic of Egger's test is indifferent to which end of the funnel is doing the tilting.

Krefeld: the most precise study is also nearly the most extreme Rothamsted, +1.29 Egger line pooled zero 0 0.005 0.010 0.014 −8 −4 0 +4 estimated change, percent per year standard error of ρ (more precise at top)
Figure 3. Funnel plot. Vertical axis is the standard error of \(\rho\), increasing downwards, so precise studies sit near the top. The shaded triangle is the pseudo 95 percent region around the pooled estimate. The dashed diagonal is the Egger regression, whose intercept is +6.87 (SE 2.18, \(t = 3.16\), \(p = 0.012\)); extrapolated to infinite precision it implies a rate of −6.93 percent per year. The asymmetry is real and it is being driven almost entirely by one point: the Krefeld study, which is both the most precise and among the steepest. That is the reverse of the usual small-study-bias signature, and it is a good reason not to read this test as evidence of publication bias.

The asymmetry here carries a sign publication bias does not predict. In the classical pattern the imprecise studies at the bottom of the funnel are the ones shifted towards large effects, because a small study with a small effect never got published. Ours does the reverse. The most extreme effect belongs to the most precise study, sitting alone at the top and well outside the pseudo-confidence region, while the bottom of the funnel holds a positive result at +1.29 percent per year, exactly the finding a publication-bias story says should be missing.

Our reading of that point runs differently. The Krefeld standard error, however correctly computed inside its own model, describes uncertainty about German nature-reserve Malaise catches rather than uncertainty about insect trends anywhere else. Its precision is real, and it is local, and a meta-analysis has no way of telling the two apart from a printed standard error. On a funnel plot that reads as an outlier. In the world it reads as a well-measured place.

We are also obliged to say that Egger's test at \(k = 11\) has very little power, and that its recommended minimum is usually quoted as ten studies, which puts us barely over the line. We ran it because we had said in advance that we would run it, which is the only defensible reason for running a test whose result you already suspect you will disbelieve. We are not leaning on it.

Take One Out, Then Take Out The Duplicates

Leave-one-out is the simplest sensitivity analysis there is: drop each study in turn, re-pool the remaining ten, and see how much the answer moves. If one study is carrying the whole result, an omission that shifts the pool a long way will say so immediately.

Nothing is. Across all eleven omissions the pooled estimate ranges from −3.05 to −2.29 percent per year, and every one of those values sits inside the confidence interval of the full-data estimate. Dropping the Krefeld study, the one carrying half the fixed-effect weight, moves the pooled figure from −2.68 to −2.29 and drops \(I^2\) from 96.3 to 87.0 percent. Substantial, and not decisive. Figure 4 plots all eleven omissions against the full-data interval.

STUDY OMITTED pooled %/yr [95% CI] omit Hallmann 2017 −2.29 [−3.44, −1.12] omit Shortall Hereford −2.53 [−4.31, −0.72] omit Shortall Rothamsted −3.05 [−4.70, −1.38] omit Shortall Starcross −2.90 [−4.59, −1.19] omit Shortall Wye −2.92 [−4.59, −1.22] omit Kaaistoep moths −2.55 [−4.39, −0.68] omit Kaaistoep beetles −2.49 [−4.26, −0.68] omit Wijster carabids −2.51 [−4.35, −0.64] omit Wepprich Ohio −2.75 [−4.55, −0.92] omit Edwards US −2.82 [−4.56, −1.05] omit Weiss NE Germany −2.64 [−4.40, −0.85] −5.5 −4.5 −3.5 −2.5 −1.5 −0.5 0 pooled estimate with that study omitted, percent per year solid line = full eleven-study pool, −2.68 %/yr shaded band = its 95% confidence interval
Figure 4. Leave-one-out. Each row is the random-effects pooled estimate with one study removed, with its 95 percent interval. The vertical line is the full-data estimate of −2.68 percent per year and the shaded band is its confidence interval. Every omission lands inside that band. The two largest movements are in opposite directions and are marked in the darker value: dropping the steepest study pulls the pool to −2.29, dropping the only positive study pushes it to −3.05.

Then the dependence problem, which has two parts. Four of our eleven effects come from one paper, three from another. The pool is not eleven independent pieces of evidence. Worse: the 2025 national butterfly compilation [11] aggregates 35 monitoring programmes across the United States, and the Ohio network analysed separately in 2019 [10] is very likely among them, which makes those two rows partly the same data counted twice. Nothing printed in either paper lets us quantify how much of the American data appears twice. So we flag it. Nine independent datasets at best, not eleven. Collapsing to one effect per paper leaves \(k = 6\) and gives −3.47 percent per year (95% CI −5.65 to −1.41), while collapsing to one effect per locality leaves \(k = 10\) and gives −2.49. Splitting by metric is more interesting. The five biomass series pool to −2.04 percent per year with an interval of −5.53 to +1.42, which crosses zero; the six abundance series pool to −3.14 with an interval of −4.39 to −1.99, which does not. Splitting by continent gives −2.92 for the nine European effects and −1.67 for the two North American ones, a gap that echoes the published finding that continental origin matters a great deal [2] and that rests here on exactly two studies. Treat it as a rhyme rather than a result.

The honest summary of this section: the pooled estimate is stable against every reasonable rearrangement of the same eleven numbers, and that stability tells you nothing at all about whether the eleven numbers are the right eleven.

The Best Argument Against This Article

Suppose the entire literature is an artefact. The case for that is better than it sounds.

Long-term monitoring sites are not chosen at random. Somebody chose them, one site at a time, for reasons that had little to do with sampling theory and a great deal to do with where insects were already known to be. A nature reserve gets designated because it is notably good. A volunteer starts a butterfly transect on the meadow full of butterflies. Funding follows the site where somebody already found a result. A university puts a field station where something is worth studying, then keeps it there for forty years, by which time the station itself is half the reason the place still looks interesting. Selection on abundance is baked into the process by which ecological monitoring comes to exist at all, and it operates on the abundance observed at the moment of selection, which carries a year-specific fluctuation sitting on top of the site's permanent quality.

Regression to the mean does the rest, manufacturing a downward trend out of nothing. If a site was picked partly because year one happened to be a good year, year two should come in worse, purely because the lucky part of year one does not repeat. Fit a straight line through a series that starts on a fluke. You will fit a decline. Nothing out in the world has to change for that line to slope downwards. The critical literature has made this case carefully and repeatedly [18], and the choice of reference year is precisely the point on which the 166-survey synthesis was challenged [4] and defended [5].

So we simulated it, deliberately setting the model up to be unkind to ourselves. Four hundred sites. Twenty-seven years, matching the Krefeld window. Permanent site quality drawn from a normal distribution, standard deviation 1.0 on the log scale. Independent year-to-year noise on top. The true trend is exactly zero. Rank the sites on their first year's observed count, monitor only the top fraction, fit an ordinary least-squares log-linear trend to each surviving series, average the slopes. Figure 5 is the result, at three levels of year-to-year noise.

our pooled estimate, −2.68 %/yr 166-survey estimate, −1.11 %/yr year noise SD = 1.0 SD = 0.6 SD = 0.2 0 −1 −2 −3 100% 50% 25% 10% 5% 2.5% fraction of sites kept, ranked on first-year observed abundance apparent trend, %/yr (true trend = 0)
Figure 5. How much decline site selection can manufacture out of nothing. Four hundred simulated sites over 27 years with a true trend of exactly zero, ranked on their first-year observed count, with only the best fraction retained for monitoring. Each point is averaged over 24 independent runs. Even under the harshest setting plotted, keeping only the top 2.5 percent of sites with year-to-year noise of SD 1.0 on the log scale, the manufactured decline reaches −1.36 percent per year. That is above the 166-survey terrestrial estimate and half of our pooled estimate, and it requires assumptions nobody would defend.

The mechanism is real and the figure shows it working. At realistic settings it is also far too small to do the job. With year-to-year noise of SD 0.6, roughly a factor of 1.8 between a good year and a bad one, and keeping the top tenth of sites, selection manufactures −0.39 percent per year out of a world with no trend in it. Call that 14.6 percent of our pooled estimate. Push the noise to SD 1.0, keep only the top 2.5 percent, an absurd degree of selection, and you reach −1.36 percent per year, which is the largest artefact this mechanism will produce under any setting anyone has ever proposed out loud. Still only half.

Three further things the simulation shows, one of which cuts against us. Turning the noise down to SD 0.2 nearly abolishes the effect, which confirms regression to the mean as the mechanism rather than some artefact of the simulation we had failed to notice. Adding a real decline of 1 percent per year underneath does not switch selection off; the observed trend becomes −1.39 percent per year at top-tenth selection, so selection inflates true declines as well as inventing false ones. And the effect rises monotonically with selection strength, so anyone arguing that monitoring programmes are sited more selectively than our worst setting holds a coherent position. Just not an evidenced one.

Determination · the objection

Site selection bias is real. Its mechanism is regression to the mean, and it submits to measurement. Under a model chosen to be unfavourable to us it accounts for roughly 15 percent of the pooled decline, rising to about 50 percent only under settings nobody can defend.

It does not manufacture 76 percent over 27 years. It does inflate whatever real signal exists, which means published rates should be read as upper bounds on the true rate at a randomly chosen place, and it gives one more reason to trust the prediction interval over the point estimate.

The Measurement Problem, Which Is The Actual Story

Set the pooled number aside.

Eleven estimates. One Malaise trap network in German reserves says 6.11 percent per year. One suction trap in Hertfordshire, thirty years long and run by a national institute, says insects went up. These are not competing claims about one quantity, with one of them mistaken. They are accurate reports of what happened in two places, made with two instruments that catch different things, over overlapping but not identical decades, summarised by two statistical models with different ideas about what a trend is.

Consider what the word "insects" is doing in a sentence like "insects are declining at 9 percent per decade." Perhaps five and a half million insect species exist, of which around a million are described. A Malaise trap samples flying insects that fly into mesh, weighted by body mass, so a good year for St Mark's flies swamps everything else in the biomass index [8]. A light trap samples moths and beetles that are phototactic, which is a behaviour, and one that artificial lighting in the surrounding countryside has been altering across the whole period of every one of these time series. A pitfall trap samples ground-active arthropods weighted by how much they walk, so a warm dry summer raises the catch without raising the population by a single beetle. A butterfly transect samples what a volunteer can name on the wing.

Four instruments. Four different functions of the underlying community, each entangled with weather, each entangled with observer behaviour, and not one of them measuring abundance in the sense a reader assumes when they meet the word in a headline. The literature knows this perfectly well; the review that lists seven challenges to interpreting insect declines puts measurement and sampling near the top of the list [16].

The consequence deserves plainer language than it usually gets. When two insect trend studies disagree, start with the instruments. Only afterwards should anyone suppose that one of them is wrong, or that the world itself differs between the two sites, and in this literature the second supposition gets made far too early and far too often. Our \(I^2\) of 96.3 percent gets read as "the sites genuinely differ." Part of it certainly is exactly that, and the Shortall traps are the proof. Some unknown and currently unknowable share of it is instrument, model choice, reference year and the decision about which taxa to count, and no meta-analysis of published summaries can separate those, because the separation would need the raw catches and one common analysis pipeline applied to every one of them.

Which is exactly what the 166-survey synthesis tried to do [2], which is why its authors got the smallest and most defensible number in the whole field, and which is also why they got a technical comment arguing their inclusion criteria and their treatment of the first year were wrong [4], and a response defending both [5]. That exchange is more informative than either paper alone. Two careful groups, the same 166 datasets, a disagreement about when the clock starts. The disagreement was not about insects.

The measurement problem does not block the science. The measurement problem is the science, at its present frontier.

None of which means the declines are imaginary, and we want to be read as saying the opposite. Ten localities, four countries, two metrics and four sampling methods, and the pooled sign is negative, the leave-one-out is stable, and the one positive result in the drawer is not statistically distinguishable from no change. Something is going down, and it is going down in more places than not. What nobody can currently tell you, from this literature, is how fast, over what area, for which insects, or with what confidence, and the confident single figures in circulation are borrowing precision that the underlying measurements do not have.

What We Would Actually Say To A Newspaper

Six sentences, in the order we would say them, with the interval attached to the estimate every time.

  1. Insect abundance and biomass are declining across most long-term monitoring sites that have been studied, and the direction of that finding is not in serious doubt.
  2. The best single number we can extract from eleven published estimates is a fall of about 2.7 percent a year, which compounds to roughly a quarter lost per decade, and it has a confidence interval from 1 to 4.3 percent a year that you should quote alongside it or not quote it at all.
  3. These studies disagree with each other far more than they disagree with zero, \(I^2 = 96.3\) percent, and the interval within which the next study is expected to land runs from an 8.1 percent annual collapse to a 3.1 percent annual increase.
  4. Freshwater insects in Europe and North America have been going the other way for decades, partly because water got cleaner [2, 18], and any account of insect decline that leaves this out is selling something.
  5. Site selection is correct in mechanism, and under an unkind model it accounts for around 15 percent of the apparent decline, so published rates are better read as upper bounds than as estimates.
  6. We are nine students with a spreadsheet and a copy of Python, we could verify eleven effect sizes and no more, and if you are going to print any of this then print the interval and not the point.
Determination · final

Pooled random-effects estimate −2.68 %/yr (95% CI −4.34 to −0.99), k = 11, from 6 papers and 10 monitoring datasets, of which two share data. I² = 96.3%, τ = 0.028, prediction interval −8.14 to +3.10 %/yr.

Confidence in the sign: high. Confidence in the magnitude: low, and deliberately so. A school club extracted this from printed abstracts and tables; it is not a systematic review, and the pooled point estimate should be read as one more entry in the drawer rather than as a summary of the field.

The calculation lives in insect-decline-meta.py, which uses nothing outside the Python standard library so that anyone can run it, and which prints its own validation checks before it prints a single result. The interactive model lets you throw studies out of the pool by clicking them, watch \(I^2\) move, and turn the selection-bias dial yourself, which is a faster route into this article than reading it twice. Think our pooled number too high? Remove the studies you object to. Then watch the heterogeneity.

It stays above 85 percent no matter which studies you throw out, which is a more durable result than anything else in this article. That is the finding.

References

  1. Hallmann, C. A., Sorg, M., Jongejans, E., Siepel, H., Hofland, N., Schwan, H., Stenmans, W., Müller, A., Sumser, H., Hörren, T., Goulson, D. & de Kroon, H. (2017). More than 75 percent decline over 27 years in total flying insect biomass in protected areas. PLoS ONE 12(10), e0185809. doi:10.1371/journal.pone.0185809
  2. van Klink, R., Bowler, D. E., Gongalsky, K. B., Swengel, A. B., Gentile, A. & Chase, J. M. (2020). Meta-analysis reveals declines in terrestrial but increases in freshwater insect abundances. Science 368, 417–420. doi:10.1126/science.aax9931
  3. van Klink, R., Bowler, D. E., Gongalsky, K. B., Swengel, A. B., Gentile, A. & Chase, J. M. (2020). Erratum for the Report “Meta-analysis reveals declines in terrestrial but increases in freshwater insect abundances”. Science 370, eabf1915. doi:10.1126/science.abf1915 [corrects the terrestrial estimate to −1.11% per year]
  4. Desquilbet, M., Gaume, L., Grippa, M., Céréghino, R., Humbert, J.-F., Bonmatin, J.-M., Cornillon, P.-A., Maes, D., Van Dyck, H. & Goulson, D. (2020). Comment on “Meta-analysis reveals declines in terrestrial but increases in freshwater insect abundances”. Science 370, eabd8947. doi:10.1126/science.abd8947
  5. van Klink, R., Bowler, D. E., Gongalsky, K. B., Swengel, A. B., Gentile, A. & Chase, J. M. (2020). Response to Comment on “Meta-analysis reveals declines in terrestrial but increases in freshwater insect abundances”. Science 370, eabe0760. doi:10.1126/science.abe0760
  6. Crossley, M. S., Meier, A. R., Baldwin, E. M., Berry, L. L., Crenshaw, L. C., Hartman, G. L., Lagos-Kutz, D., Nichols, D. H., Patel, K., Varriano, S., Snyder, W. E. & Moran, M. D. (2020). No net insect abundance and diversity declines across US Long Term Ecological Research sites. Nature Ecology & Evolution 4, 1368–1376. doi:10.1038/s41559-020-1269-4
  7. Welti, E. A. R., Joern, A., Ellison, A. M., Lightfoot, D. C., Record, S., Rodenhouse, N., Stanley, E. H. & Kaspari, M. (2021). Studies of insect temporal trends must account for the complex sampling histories inherent to many long-term monitoring efforts. Nature Ecology & Evolution 5, 589–591. doi:10.1038/s41559-021-01424-0
  8. Shortall, C. R., Moore, A., Smith, E., Hall, M. J., Woiwod, I. P. & Harrington, R. (2009). Long-term changes in the abundance of flying insects. Insect Conservation and Diversity 2, 251–260. doi:10.1111/j.1752-4598.2009.00062.x
  9. Hallmann, C. A., Zeegers, T., van Klink, R., Vermeulen, R., van Wielink, P., Spijkers, H., van Deijk, J., van Steenis, W. & Jongejans, E. (2020). Declining abundance of beetles, moths and caddisflies in the Netherlands. Insect Conservation and Diversity 13, 127–139. doi:10.1111/icad.12377
  10. Wepprich, T., Adrion, J. R., Ries, L., Wiedmann, J. & Haddad, N. M. (2019). Butterfly abundance declines over 20 years of systematic monitoring in Ohio, USA. PLoS ONE 14(7), e0216270. doi:10.1371/journal.pone.0216270
  11. Edwards, C. B., Zipkin, E. F., Henry, E. H., Haddad, N. M., Forister, M. L. et al. (2025). Rapid butterfly declines across the United States during the 21st century. Science 387, 1090–1094. doi:10.1126/science.adp4671
  12. Weiss, F., von Wehrden, H. & Linde, A. (2024). Long-term drought triggers severe declines in carabid beetles in a temperate forest. Ecography 2024, e07020. doi:10.1111/ecog.07020
  13. Lister, B. C. & Garcia, A. (2018). Climate-driven declines in arthropod abundance restructure a rainforest food web. Proceedings of the National Academy of Sciences 115, E10397–E10406. doi:10.1073/pnas.1722477115
  14. Schowalter, T. D., Pandey, M., Presley, S. J., Willig, M. R. & Zimmerman, J. K. (2021). Arthropods are not declining but are responsive to disturbance in the Luquillo Experimental Forest, Puerto Rico. Proceedings of the National Academy of Sciences 118, e2002556117. doi:10.1073/pnas.2002556117
  15. Macgregor, C. J., Williams, J. H., Bell, J. R. & Thomas, C. D. (2019). Moth biomass has fluctuated over 50 years in Britain but lacks a clear trend. Nature Ecology & Evolution 3, 1645–1649. doi:10.1038/s41559-019-1028-6 [see Author Correction, Nature Ecology & Evolution 5, 865–883 (2021), doi:10.1038/s41559-021-01449-5]
  16. Didham, R. K., Basset, Y., Collins, C. M., Leather, S. R., Littlewood, N. A., Menz, M. H. M., Müller, J., Packer, L., Saunders, M. E., Schönrogge, K., Stewart, A. J. A., Yanoviak, S. P. & Hassall, C. (2020). Interpreting insect declines: seven challenges and a way forward. Insect Conservation and Diversity 13, 103–114. doi:10.1111/icad.12408
  17. Møller, A. P. (2019). Parallel declines in abundance of insects and insectivorous birds in Denmark over 22 years. Ecology and Evolution 9, 6581–6587. doi:10.1002/ece3.5236 [RETRACTED 2026; retraction notice doi:10.1002/ece3.73864]
  18. Haase, P., Bowler, D. E., Baker, N. J., Bonada, N., Domisch, S. et al. (2023). The recovery of European freshwater biodiversity has come to a halt. Nature 620, 582–588. doi:10.1038/s41586-023-06400-1
  19. Sockman, K. W. (2025). Long-term decline in montane insects under warming summers. Ecology 106, e70187. doi:10.1002/ecy.70187
  20. DerSimonian, R. & Laird, N. (1986). Meta-analysis in clinical trials. Controlled Clinical Trials 7, 177–188. doi:10.1016/0197-2456(86)90046-2
  21. Higgins, J. P. T. & Thompson, S. G. (2002). Quantifying heterogeneity in a meta-analysis. Statistics in Medicine 21, 1539–1558. doi:10.1002/sim.1186
  22. Egger, M., Davey Smith, G., Schneider, M. & Minder, C. (1997). Bias in meta-analysis detected by a simple, graphical test. BMJ 315, 629–634. doi:10.1136/bmj.315.7109.629