Science Journaling Club Founded 2024

REVIEW · SYSTEMATIC REVIEW · CONSERVATION METHODS

Can a Cup of Pond Water Tell You What Lives in It? Reviewing Environmental DNA Detection Rates

Written jointly by the Science Journaling Club

Review · Peer-edited by the club review board · LaTeX source · Our calculation · Interactive model

Abstract Filter some water, lyse what the membrane caught, run a PCR, and you are supposed to learn which animals were recently swimming there. We went looking for every study we could open that surveyed the same water bodies twice, once with a filter and once with a net or a torch or a camera, and that reported a detection probability for both. We found twelve we could pool and seven more we could not. On a random effects model the odds that one eDNA survey unit detects an occupied site are 2.05 times the odds for one conventional survey unit, with a 95% interval of 0.95 to 4.41, which does not exclude a tie. The heterogeneity is the actual result: I² = 80.5%, and the interval predicting the next study runs from an odds ratio of 0.15 to 27. Six of the twelve contrasts favour the filter and six favour the net. The strongest single pattern is taxonomic: amphibians sit at an odds ratio near four, reptiles below one, and no protocol tuning we can see closes that gap. Then we set the pooling aside and look at the arithmetic nobody can get around: with a per-sample false positive rate of 0.02 and a true occupancy of 0.02, a single positive is real only 38% of the time. Requiring two positives out of two lifts that to 0.95 and drops sensitivity from 0.60 to 0.36, and that trade is the method rather than a small print item.
1 · SampleCollect water. Anything from 90 mL to 60 L.Fails when: the DNA is not where your bottle is.
2 · FilterPass it through 0.45 to 1.5 µm.Fails when: the filter clogs and you quietly shrink the volume.
3 · ExtractLyse the filter, bind, wash, elute.Fails when: humic acid rides along and inhibits.
4 · AmplifyqPCR or metabarcoding primers.Fails when: primers hit a congener, or a tag jumps.
5 · DetectCall the site positive, or not.Fails when: the call is treated as truth.

Sample failure: the animal is fifty metres upstream

The question we actually wanted answered

We want to use this, and that is not a neutral place to start a review.

Start there, because the wanting shapes everything below. Several of us would like to take a bottle to the pond behind the sports field, filter a litre of water through a glass fibre membrane, and find out whether the great crested newts that somebody saw in 2019 are still breeding there, instead of spending four nights in April holding a torch over the same water and counting eyeshine. Environmental DNA, eDNA to everyone who handles it, means the loose genetic material animals shed into water: skin cells, mucus, gametes, faeces, all of it drifting downstream and degrading on a clock nobody has calibrated. Catch it on a filter, ask the filter what swam past, and you have a promise that is now about fifteen years old and still not fully cashed.

Reading around, we could not find a straight answer to the operational question. Not "does eDNA work", which is too vague to be false. The meaner version: per unit of effort, is the bottle better than the net? One water sample against one seine haul, one filter against one trap-night. That ratio is the exchange rate you need before you can plan a season of fieldwork, and it decides whether the method saves you anything at all or just moves the cost from the pond to the lab.

So we did a review. Twelve paired comparisons pooled, seven more studies found and set aside with the reason written down, and a random effects model built from scratch so we could watch every step of it. Nobody got the clean answer. The messy one turned out to be better, and the best part of the mess is a failure mode that occupancy statisticians have been shouting about for a decade while field protocols carry on ignoring it.

Filter failure: clogging, and the volume you meant to take

What the protocol actually does, stage by stage

Before any statistics, walk the physical chain from bottle to call. Five stages, each with its own way of losing the signal, and the twelve studies disagree very largely because they made different choices at every one of them.

Collecting. The spread here is enormous and nobody talks about it enough. Biggs and colleagues, working on great crested newts across 35 UK ponds, took twenty 30 mL scoops from around the pond edge, pooled them into 600 mL, then subsampled six tubes of 15 mL into ethanol and sodium acetate and precipitated the DNA out [8]. Ninety millilitres of water actually reach the extraction, which is roughly a wine glass of water standing in for a pond you could not throw a stone across. At the other end, Lopes and colleagues pumped 20 L or 60 L through a capsule in Brazilian streams [21], and Lyet and colleagues went to 30 to 80 L per sample in a savanna catchment. Between 90 mL and 60 L sits a factor of about 670. Nobody would tolerate a 670-fold spread in the effort of a conventional survey and then call two such surveys comparable units, yet the eDNA literature does it every issue.

Filtering. Pore size across our twelve runs 0.45 µm to 1.5 µm, with glass fibre and cellulose nitrate both represented and nobody reporting what the choice cost them. A 0.45 µm membrane catches more and clogs sooner, which matters in a way the papers rarely quantify. Once the membrane blinds off, the field team takes less water than the protocol says; sometimes they swap to a coarser pore, sometimes they open a second filter and carry on. Moss and colleagues say so outright, listing smaller volumes and a third filter among the things they did when their membranes clogged mid-sample [14]. So the volume you report is partly a readout of how much silt was suspended in the pond that morning, and volume and turbidity arrive at the freezer already tangled.

Extracting. Our thinnest section, because the papers mostly report a kit name and stop. Inhibition is the failure here: humic acids, mostly, leached out of leaf litter, riding through the spin column into the eluate and sitting on the polymerase. Moss tracked which samples needed inhibitor removal and carried it as a covariate [14].

Amplifying. Two families. A species-specific qPCR assay aims one primer pair, usually with a hydrolysis probe, at a single animal, and hands back a cycle threshold you can argue about later. Metabarcoding does the opposite. Degenerate primers bind conserved flanking regions across a whole group, amplify whatever sits between them, and the sequencer sorts out the rest afterwards, which is why a metabarcoding run answers questions you never asked. Most European pond work inherits Valentini's batra and teleo pairs [7], written one for amphibians and one for bony fish, and inherits their blind spots with them. Specificity is a design property rather than a fact of nature, and a primer pair that cannot separate your target from its congener will report the congener as your target, cheerfully, forever.

Detecting. Somebody decides. One positive PCR replicate out of three, or two out of twelve. A read count over some threshold somebody picked. The interesting statistics live here, and section 10 is about nothing else.

Detect failure: calling a non-detection an absence

Borrowing the occupancy machinery we already built

The club has been here before, and every number below sits on the same frame. In Absence of Evidence we worked through the occupancy model of MacKenzie and colleagues [1], and nothing in this review improves on it or needs to. One paragraph of restatement, then.

A site either holds the species or it does not; call that probability \(\psi\). One survey at an occupied site detects with probability \(p\), and misses with probability \(1-p\), and those are the only two outcomes the model allows itself. Run \(K\) independent surveys there and the chance of seeing it at least once is

$$p^{*} = 1 - (1-p)^{K}$$

which is the only equation you need to keep the rest of this straight. One consequence wrecks a lot of naive comparisons. A method with a low \(p\) can be dragged up to look excellent simply by running it more often, because \(p^*\) climbs fast and the arithmetic does not care how the visits were paid for. Valentini's amphibian result says exactly this: a single traditional visit has \(p = 0.58\), so four visits give \(1 - 0.42^4 = 0.969\), which matches what one eDNA sample achieved on its own [7]. Those are not the same statement about the method; they are the same statement about a budget.

So the effect size is a per-unit detection probability, logit-transformed and differenced between the two methods, one row per study and no study counted twice anywhere. Writing \(p_{e}\) for eDNA and \(p_{c}\) for the conventional comparator,

$$y = \operatorname{logit}(p_{e}) - \operatorname{logit}(p_{c}) = \log\!\left(\frac{p_{e}/(1-p_{e})}{p_{c}/(1-p_{c})}\right)$$

a log odds ratio, with \(y > 0\) meaning the filter beats the net per unit of effort. The logit transform earns its place here. Detection probabilities pile up against 1.0, where a difference of 0.04 between 0.95 and 0.99 is a much larger statement than the same 0.04 between 0.50 and 0.54, and the log odds scale knows that while a raw difference does not.

Sample failure: the literature you can open is not the literature

How the twelve were chosen, and how the code was tested

Rules first, fixed before we looked at any pooled number.

A study was eligible if it surveyed water bodies with both an eDNA protocol and a conventional method, at the same sites, and reported something from which a per-unit detection probability could be recovered for each. Terrestrial and airborne eDNA work was out even where the numbers were beautiful, which cost us a spotted lanternfly study reporting 0.77 per sample against 0.25 per trap. Studies reporting only how many species each method found were out, because species richness is a different quantity from a detection probability and no arithmetic converts one into the other. Studies where the conventional method returned a flat zero were out, because the odds ratio is not finite and a continuity correction there would be an invention rather than an estimate. Studies reporting ranges instead of estimates were out too, which cost us Smart and colleagues on an invading smooth newt population in Melbourne, whose per-sample eDNA detection is given as 0.29 to 1.0 across sites against 0.01 to 0.26 per bottle trap [19]; a range is not a number with a standard error on it.

Four studies reported several target species. We needed one row each, and picking the flattering species would have been trivially easy, so we fixed a rule before extraction: take the median contrast by log odds ratio. For Hinlo's three invasive fish that picks redfin perch, between common carp at −1.17 and Oriental weatherloach at +4.36 [10]; for Moss's amphibians it picks the California newt, between the red-legged frog at +3.24 and the bullfrog at −0.18 [14]. The rule is not clever, but deciding it in advance is the only property that matters.

A published 95% interval converts straight to a standard error on the logit scale; a standard error on the probability scale needs the delta method first. Three papers gave a point estimate and nothing else; for those we reconstructed a binomial standard error from the number of independent survey units reported, \(\mathrm{SE} = \sqrt{1/(n\,p\,(1-p))}\), and flagged the row. The reconstruction is conservative, since it ignores the replicate subsamples inside each unit, and it remains a reconstruction rather than the authors' own uncertainty, which is a real difference.

Eleven of the twelve numbers we read in the paper itself. The twelfth, Schmelzle and Kinziger on tidewater goby, we could not open; its 0.74 against 0.39 is quoted consistently in several secondary sources and in the paper's own abstract, but we never saw the table it came from. The table below marks it, and a sceptical reader should pull that row first.

Checking the machinery before trusting it

Pooling code is easy to write and easy to get subtly wrong, so before running it on anything real we made it prove two things.

First, recovery. We generated fourteen synthetic studies from a world with a known true mean log odds ratio of 0.800 and a known between-study standard deviation of 0.450, gave each a different within-study variance, and pooled them. Over 4,000 repeats the mean pooled estimate came back at 0.8017, a bias of +0.0017, and the mean \(\hat{\tau}^{2}\) at 0.2045 against a true 0.2025, which is as close as a method-of-moments estimator gets. The intervals covered the true value 91.8% of the time against a nominal 95%. That undercoverage is not a bug in our code; it is a known property of the DerSimonian-Laird interval with few studies, which treats \(\hat{\tau}^{2}\) as though it were known. So our own 95% interval is too narrow, and the real one runs wider than 0.95 to 4.41.

Second, the identity. When \(Q \le \mathrm{df}\) the estimator truncates \(\hat{\tau}^{2}\) to zero, at which point the random effects weights \(1/(v_i + \hat{\tau}^{2})\) become the fixed-effect weights \(1/v_i\) exactly, and the two estimates must be identical rather than merely close. We built a seven-study set with no real heterogeneity, got \(Q = 0.0082\) on 6 degrees of freedom and \(\hat{\tau}^{2} = 0\), and the two estimates agreed to 0.000e+00 on both the point estimate and its standard error. Not approximately. Bitwise.

Third, and less formally, we fed Egger's regression a funnel that was symmetric by construction and confirmed the intercept came back statistically indistinguishable from zero (p = 0.23). All three checks print at the head of the output file, before any real data is touched, which is exactly where a check belongs and almost never where one appears.

Extract failure: the number you want is in a figure, not a table

The table, and the arithmetic on it

Twelve rows, one study each, ten columns. One target, one pair of detection probabilities, one log odds ratio.

StudyTargetComparator UnitWaterp(eDNA)p(conv) log ORSEWhere the number is
Biggs 2015 [8]Great crested newtTorch countvisit0.09 L0.990.75+3.440.84Results text: 139 of 140 visits; torch 75%, traps 76%, eggs 44%
Valentini 2016 [7]Amphibians, 39 French pondsOne traditional visitvisit0.09 L0.970.58+3.150.63Results: "0.97 (CI = 0.90–0.99) vs. 0.58 (CI = 0.50–0.63)"
Schmelzle 2016 [9]Tidewater gobySeine haulsamplen/r0.740.39+1.490.57Abstract only; SE reconstructed from n = 29 sites
Hinlo 2017 [10]Redfin perchFyke net setsite-season12 L0.210.00+2.181.56Table 1 detection matrix, 14 paired site-seasons
Eiler 2018 [11]Pool frogCalling and visual countvisit0.85 L0.380.40−0.080.41Results: "0.38 per sample … compared to 0.40"; SE from n = 49
Rose 2019 [12]Northern watersnakeOne trap-night, 30 trapssample0.5 L0.440.74−1.291.06Reported: eDNA 0.44 (0.15–0.75), trap 0.74 (0.48–0.95)
Akre 2019 [13]Wood turtleOne-hour visual surveysample2 L0.550.88−1.790.98Reported: eDNA 0.55 (0.38–0.71), VES 0.88 (0.58–0.98)
Moss 2022 [14]California newtSeine haulvisit0.25 L0.650.69−0.180.62Results: MB 0.65 (0.44–0.81), seine 0.69 (0.48–0.84)
Quilumbaquin 2023 [15]Amphibians, Ecuadorian AmazonVisual encounter surveyvisit1 L0.420.17+1.260.12Abstract: "DP of 0.42 CI [0.40–0.45] … DP of 0.17 CI [0.14–0.20]"
Li 2024 [16]Amphibians, Zhoushan islandsLine transectsite2 L0.540.24+1.310.67Results: "0.54 … 0.24"; no interval given, SE from n = 21 islands
Di Girolamo 2024 [17]American minkOne camera-trap weekvisit0.5 L0.250.36−0.520.82Results: eDNA ρ = 0.25 (SE 0.08), camera ρ = 0.36 (SE 0.16)
Dougherty 2025 [18]AlewifePurse seine surveysite-season3.55 L0.571.00−2.461.63Results: 3 false negatives out of 7 lakes; seining missed none

Six positive, six negative. Unweighted mean of the twelve log odds ratios: +0.53, median +0.59, and the sign split exactly even. Largest +3.44, smallest −2.46, a spread of 5.90 log odds, which on the odds scale is a factor of 365 between the best case and the worst. Standard deviation of the twelve: 2.05. Six of the twelve standard errors exceed 0.8. One row on its own carries almost nothing.

eDNA detection probability across the twelve runs 0.21 to 0.99, median 0.53; the conventional comparators run 0.00 to 1.00, median 0.48. Both methods therefore cover almost the entire available range, which already tells you that no single number describes either of them well enough to plan a survey around. Figure 3 draws each pair as a line.

Detection probability of one survey unit 0.00 0.00 0.25 0.25 0.50 0.50 0.75 0.75 1.00 1.00 conventional eDNA Biggs 2015 Valentini 2016 Schmelzle 2016 Moss 2022 Dougherty 2025 Akre 2019 Li 2024 Rose 2019 Quilumbaquin 2023 Eiler 2018 Di Girolamo 2024 Hinlo 2017 up-sloping lines: eDNA ahead: 6 of 12 down-sloping: conventional: 6
Figure 3. Each study’s pair of per-unit detection probabilities, conventional method on the left axis and eDNA on the right. Solid up-sloping lines are the six contrasts where the filter beat the net; dashed down-sloping lines are the six where it did not. The vertical spread on both axes is the point: detection probability for a single eDNA sample ranges from 0.21 to 0.99 across these studies, which is most of the available range.

Amplify failure: pooling things that were never the same measurement

Pooling it, and the result nobody wanted

We used a DerSimonian and Laird random effects model [4], which assumes each study estimates its own true effect drawn from a distribution with mean \(\mu\) and between-study variance \(\tau^{2}\), and estimates \(\tau^{2}\) from the observed spread using the method-of-moments formula. All of it sits in the linked Python, about forty lines with no library in the way, so nothing in the pooling happens where you cannot watch it.

The pooled log odds ratio comes to +0.718. On the odds scale: 2.05, with a 95% interval of 0.95 to 4.41. So the filter roughly doubles the odds of detection per unit of effort, on average, and the interval still contains a tie between the two methods. A two-sided p value of 0.066. Figure 1 draws all twelve contrasts with the pooled diamond beneath them.

Study log odds ratio (eDNA vs conventional) OR [95% CI] -4 -3 -2 -1 0 1 2 3 4 conventional better eDNA better Biggs 2015 31.29 [5.99, 163.5] Valentini 2016 23.41 [6.86, 80.0] Schmelzle 2016 4.45 [1.46, 13.6] Hinlo 2017 8.83 [0.41, 188.7] Eiler 2018 0.92 [0.41, 2.1] Rose 2019 0.28 [0.03, 2.2] Akre 2019 0.17 [0.02, 1.1] Moss 2022 0.83 [0.25, 2.8] Quilumbaquin 2023 3.54 [2.79, 4.5] Li 2024 3.72 [0.99, 13.9] Di Girolamo 2024 0.59 [0.12, 2.9] Dougherty 2025 0.09 [0.00, 2.1] Pooled (RE) 2.05 [0.95, 4.41] 95% prediction
Figure 1. The twelve paired contrasts, on the log odds scale, with eDNA to the right of the dashed null line. Box area is the random-effects weight. Arrowheads mean the interval runs past the edge of the frame. The diamond is the pooled estimate, an odds ratio of 2.05 (0.95 to 4.41); the bar beneath it is the 95% prediction interval, which runs from an odds ratio of 0.15 to 27.4 and is the honest summary of this literature. Six contrasts sit either side of zero.

Now the part that matters more than the estimate. Cochran's Q, which measures how much the studies scatter beyond what their own standard errors allow, comes to 56.3 on 11 degrees of freedom, a p value below 0.0001. The derived I² is 80.5%, so four fifths of the variation is real disagreement rather than sampling noise [5]. The between-study standard deviation \(\tau\) comes to 1.10 on the log odds scale. Against a pooled mean of 0.72, that is enormous.

Put that into a prediction interval, the interval covering where a thirteenth study's true effect would be expected to land, and it runs from an odds ratio of 0.15 to 27.4. Read it on the probability scale. If your conventional method detects an occupied site half the time, this literature predicts that the eDNA protocol you are about to use will detect it somewhere between 13% and 96% of the time. Call that a forecast and you are being generous; a shrug with a confidence interval attached, more like.

So the headline of this review is not 2.05; the headline is 80.5%.

Detect failure: the studies that never got written

Working notes from the club table

Session 1 · deciding what counts as one survey

Argument for forty minutes about whether one water sample and one seine haul are comparable units. They are obviously not. A seine haul runs ten minutes of two people's time; a bottle is ninety seconds of one person's, plus the filtering back at the car with the pump running off a leisure battery. Settled by recording the unit on every row, then testing it as a moderator. Nobody was happy. Still the right call.

Session 2 · the Hunter problem

Burmese python, Florida. eDNA detection per sample between 0.59 and 0.87. Conventional effort: three captures in 5,935 trap-nights, and zero standardised visual sightings [20]. The odds ratio is infinite. Somebody proposed a Haldane correction. Somebody else pointed out that the correction would then be entirely responsible for the answer, which ended the argument, and the row stayed out of the table. We think that decision is correct and we also think it drags our pooled estimate downward, because the excluded case is the most extreme eDNA win anywhere in this literature. Both true at once.

Session 3 · running the funnel

Egger's test regresses each study's standardised effect on its precision; a non-zero intercept means small imprecise studies sit systematically off to one side, which is the classic signature of selective publication [6]. Our intercept: −0.94, SE 0.90, t = −1.04 on 10 df, p = 0.32. Which is to say: nothing detected, and with twelve studies and this much heterogeneity the test has almost no power, so "nothing detected" means close to nothing. Reported because we said we would, and the funnel in Figure 2 is drawn for the same reason.

Funnel: precision against effect -6 -4 -2 0 2 4 6 log odds ratio 0.0 0.5 1.0 1.5 standard error Biggs ’15 Valentini ’16 Schmelzle ’16 Hinlo ’17 Eiler ’18 Rose ’19 Akre ’19 Moss ’22 Quilumbaquin ’23 Li ’24 Di Girolamo ’24 Dougherty ’25 pooled 0.72 null Egger intercept -0.94 (SE 0.90), p = 0.32
Figure 2. Funnel plot: effect against standard error, with the pseudo 95% region drawn around the pooled estimate. Under no small-study bias and no heterogeneity, points would fall inside the shaded wedge in a symmetric cloud. They do not, but that is mostly heterogeneity rather than asymmetry: Egger’s intercept is −0.94 (SE 0.90, p = 0.32). With twelve studies the test has almost no power, so this is weak evidence of nothing.

Session 4 · the thing that actually worried us

If publication bias operates here, it runs the opposite way from usual, and the funnel is the wrong instrument for finding it. A study that finds eDNA works gets written up as a method validation. A study that finds it does not work gets written up only if someone is stubborn, and Rose and colleagues, whose title begins "Traditional trapping methods outperform eDNA sampling", clearly were [12]. So the negative results in our twelve are, we suspect, survivors of a much larger set that went unwritten. That would inflate the pooled estimate, and the funnel cannot see it because the missing studies are not missing from one corner of the funnel, they are missing from a whole methodology.

Session 5 · leaving one out

Dropped each study in turn and refitted. The pooled odds ratio moves between 1.60 and 2.50. The sign never flips. I² never falls below 78%. Dropping Biggs gives 1.64, dropping Akre gives 2.50, and eight of the twelve refits have intervals containing zero. So no single study carries the result, and no single study rescues it.

Filter failure: confounding volume with everything else

What moves the number

With I² at 80.5% the interesting question stops being "what is the average" and becomes "what explains the spread". Four candidates went in with us, water volume processed, target abundance, primer specificity and season, and only two of them survive contact with what these papers actually bothered to report.

Taxon, which we did not expect to be the strongest signal

Split the twelve by broad group and the amphibian studies pool to a log odds ratio of +1.37 (interval +0.44 to +2.30), an odds ratio near four, while the two reptile studies pool to −1.56 (interval −2.97 to −0.15), an odds ratio of 0.21, with I² of exactly zero between them. The three fish studies land at +0.64, interval −1.71 to +2.98. No information at all. Two contrasts is a thin basis for anything, but the mechanism is at least plausible: a larval amphibian sheds continuously into a small closed pond, while a watersnake or a wood turtle spends most of its time out of the water entirely and contributes DNA in brief visits. The turtle and the snake do not show eDNA failing. They show an animal that barely touches the water it lives beside.

Water volume, where the between-study answer is backwards

Regress the log odds ratios on the log of the volume processed, weighting by random effects precision, and the slope is −1.42 per decade of litres, standard error 0.74, p = 0.055. Negative. More water, worse relative performance. The volumes span 0.09 L to 12 L, a factor of 133, which leaves plenty of range to work with.

Confounded, and we can name the confounder. The two studies at 0.09 L are the UK and French pond studies, which precipitate DNA out of a small pooled sample drawn from a small, still, warm, DNA-saturated pond full of breeding amphibians. The studies at 2 L and above work streams, lakes and rivers, where the DNA is dilute, moving, and further from whatever shed it in the first place. Volume is standing in for habitat. Habitat is doing the work. Real arithmetic on real numbers, answering a question nobody asked.

For the question people actually ask, you need a study that varied volume alone. Lopes and colleagues did. They put 20 L and 60 L through the capsule at the same sampling points in four Brazilian streams, which is the comparison everybody wants and almost nobody runs [21]. Per-sample detection rose from 0.614 to 0.761 for Hylodes phyllodes, from 0.570 to 0.649 for Hylodes asper, and from 0.154 to 0.596 for the scarcest of the three, Cycloramphus boraceiensis. Mean log odds ratio for tripling the volume: +1.04, an odds ratio of 2.83, and the largest gain by far goes to the rarest of the three species. Within a study, volume helps a lot. Between studies, it predicts the opposite. Figure 5 sets the two side by side. The difference between them is the most useful methodological lesson in this review.

Between studies Within one study 0.1 1 10 -3 -2 -1 +0 +1 +2 +3 +4 litres of water processed (log) log odds ratio slope -1.42 per decade (SE 0.74, p = 0.055) 0.00 0.25 0.50 0.75 1.00 20 L 60 L per-sample detection H. phyllodes H. asper C. boraceiensis mean +1.04 log odds
Figure 5. Left: the between-study meta-regression of log odds ratio on water volume, weighted by random-effects precision. The slope is −1.42 per decade of litres (SE 0.74, p = 0.055), pointing the wrong way, because the small-volume studies are warm amphibian ponds and the large-volume studies are dilute running water. Right: the same question asked inside one study, where Lopes and colleagues filtered 20 L and 60 L at the same points. Every species improves, and the rarest improves most, for a mean log odds ratio of +1.04.

Survey unit, which is the reviewer's own artefact

Nine of our twelve contrasts compare whole survey rounds against whole conventional rounds; the other three compare single samples. Split them and the round-level group pools to an odds ratio of 2.78 with an interval of 1.18 to 6.5, while the sample-level group pools to 0.66 with an interval of 0.07 to 6.60. The difference between the subgroups is not significant (z = −1.15, p = 0.25), and with three studies on one side it could not have been. But the direction is the one you would predict from \(p^{*} = 1-(1-p)^{K}\): a "round" of eDNA usually bundles several filters against a single conventional effort, so the round-level comparisons are quietly comparing more eDNA effort with less net effort. We think that is part of why the literature reads kindly to eDNA.

Amplify failure: the review's own conclusion

The strongest case against everything above

The strongest case against us runs as follows, and we have put it as forcefully as we know how.

The pooled number is meaningless and we said so ourselves. If I² is 80.5% and the prediction interval spans a factor of 180 in odds, then reporting 2.05 at all misleads, whatever caveats trail after it. Readers remember point estimates and forget intervals. The honest output of this analysis is a refusal to pool, and by pooling anyway we have manufactured a number that will be quoted without its \(\tau\).

We accept most of this. Our defence is narrow: the pooled estimate is reported here as the input to the heterogeneity statistics rather than as a recommendation, and every figure and the entire interactive companion display the spread rather than the centre. Whether that is enough is a fair thing to argue about.

Twelve is not a literature. Fediajevaite and colleagues screened 535 empirical eDNA papers and found 194 that compared against a traditional method, of which only 49 reported a quantitative detection probability for both [3]. We found twelve. We are working from roughly a quarter of the comparable set, selected by what we could open and read, which is a sampling process with no theory behind it whatsoever.

This one lands squarely. Our twelve skew open-access, skew recent, skew towards full occupancy output. An article in a paywalled fisheries journal reporting an eDNA failure is exactly the kind of thing our procedure cannot see, and we have no way to estimate how many of those there are.

The median-contrast rule throws away most of the data. Moss reports six species and we used one. A multilevel model with species nested in study would use all six, separate within-study from between-study variance, and answer the question better than we can with twelve rows. It would also be well beyond what we can implement and check by hand, which is the actual reason we did not do it, and "we could not verify the code" is a much weaker justification than "the simple approach is adequate".

The unit problem is fatal, not merely noted. If one row compares a 500 mL bottle against thirty traps set overnight and another compares a 600 mL scoop against a ten-minute torch walk, then \(\tau^{2}\) is measuring the variance of our own extraction decisions as much as anything biological. A reader could reasonably conclude that the 80.5% is our artefact.

Our answer: partly ours, and not only ours. The two reptile contrasts, both sample-level, both against genuinely intensive conventional effort, agree with each other to an I² of zero and sit at an odds ratio of 0.21. Two studies that agree with each other that precisely, while disagreeing that hard with the amphibian rows, look like a real biological split rather than a clerical one. But we cannot prove it with twelve rows, and if somebody redoes this with a proper multilevel model on all 49 of Fediajevaite's comparable studies, we would expect the pooled estimate to move and I² to stay high.

Detect failure: treating a positive as a presence

The false positive problem, which is the interesting one

Everything up to this point has treated a positive result as the truth. A positive is not the truth, and the way it fails is genuinely lovely, in the sense that a trapdoor is lovely once somebody has shown you where it is.

Write \(p_{11}\) for the probability that an occupied site produces a positive, the true positive rate, which is the \(p\) we have been pooling all along. Write \(p_{10}\) for the probability that an unoccupied site produces a positive anyway. That happens through contamination in the field kit, contamination in the lab, primers cross-amplifying a related species, DNA carried down from a population upstream, or index hopping between samples on a sequencing run. Then for a site with prior occupancy \(\psi\), one positive sample gives

$$\Pr(\text{present} \mid +) = \frac{\psi\,p_{11}}{\psi\,p_{11} + (1-\psi)\,p_{10}}$$

which is Bayes' rule wearing a lab coat, and is the standard framing in the false-positive occupancy literature [2]. Figure 4 plots it. Now look at what it does.

Set \(p_{11} = 0.60\), a middling value from our table. Set \(p_{10} = 0.02\), a contamination rate nobody would call alarming. At \(\psi = 0.50\) a positive is right 96.8% of the time. At \(\psi = 0.10\), 76.9%. At \(\psi = 0.02\), 38.0%, and the majority of your positives are now wrong. At \(\psi = 0.005\), which is what "we think this species may have recolonised one pond in the county" actually means, a positive is right 13.1% of the time.

What one positive is worth, as the species gets rarer 0.00 0.25 0.50 0.75 1.00 0.002 0.005 0.01 0.02 0.05 0.1 0.2 0.5 true occupancy of the site, psi (log scale) P(present | one positive) p10 = 0.001 p10 = 0.005 p10 = 0.02 p10 = 0.05 0.38 one positive, psi = 0.02, p10 = 0.02 two of two rule, p10 = 0.02
Figure 4. Posterior probability that a site is genuinely occupied given one positive eDNA sample, against true occupancy, with the true positive rate fixed at p₁₁ = 0.60. Each solid curve is a different false positive rate. The marked point is the working example in the text: at ψ = 0.02 and p₁₀ = 0.02 a single positive is right 38% of the time. The dashed curve is the two-of-two replicate rule at the same p₁₀, which recovers 0.95 at that occupancy and costs sensitivity, dropping the true positive rate from 0.60 to 0.36.

The structure of that collapse is what makes it dangerous. The false positive rate did not change. The assay did not get worse. Only the rarity of the animal changed. Rarity is the exact condition under which people reach for this method, so the method degrades fastest precisely where it is being sold hardest, which is a nasty property for a survey tool. Detecting a common carp in a river full of common carp is the case where a positive is trustworthy and also the case where nobody needed the test.

A fix exists, and it costs: require \(r\) independent replicates all positive before calling a detection. If replicates are independent, the true positive rate becomes \(p_{11}^{r}\) and the false positive rate \(p_{10}^{r}\), and because \(p_{10}\) is small the second falls far faster than the first. At \(\psi = 0.02\) with our numbers, one replicate gives a posterior of 0.380, two of two gives 0.948, and three of three gives 0.998. Meanwhile sensitivity drops from 0.60 to 0.36 to 0.216. You are trading false positives for false negatives at a favourable rate. No setting of \(r\) avoids the trade.

None of the twelve studies in our table estimated \(p_{10}\), though several ran field and extraction blanks and reported no amplification. A loose bound, and not the same thing as an estimate. Ficetola and colleagues showed a decade ago that with enough PCR replicates the false positive rate becomes estimable from the data itself, and the models to do it have been sitting in the literature since [2]. The gap between what the statistics can do and what the field protocols report is the largest open problem in this whole area, and it is much larger than the question of whether the filter beats the net.

Sample failure: a survey design nobody checked the arithmetic on

What we would do with a cup of pond water

Suppose we could actually do it. Suppose the permits came through and somebody lent us a peristaltic pump. The protocol below traces every item back to a number somewhere above it rather than to a preference.

Filter more water than feels necessary. The only within-study experiment we found put tripling the volume at an odds ratio of 2.83 per sample, with the largest gain going to the rarest species [21]. If the filter clogs, use a second filter rather than stopping, and write down the volume that actually went through, because that number is the one the analysis needs and it is the one most often missing.

Take more samples rather than better ones. Moss and colleagues reach this conclusion directly, and the occupancy arithmetic backs it: with \(p = 0.55\) per sample, as Akre reports for wood turtles [13], four samples give \(1 - 0.45^{4} = 0.959\). Two samples give 0.798. The fourth sample buys more than any plausible improvement to a single sample would.

Require two independent positives, and define independent at the bottle. One positive at \(\psi = 0.02\) lands wrong more often than right. Two of two from separate bottles takes you to 0.95 on the same assumptions, at a sensitivity cost you can calculate before you leave the house.

Run the conventional survey anyway, at least at first. Six of our twelve contrasts favour the conventional method, which is not a footnote. Di Girolamo and colleagues found camera traps and eDNA both around a quarter to a third per unit, and the two methods combined reaching 0.48 [17]. The methods miss different sites. Most of these papers report exactly that. Reading them as a contest is the mistake.

Report \(p_{10}\), or say that you did not. Clean blanks are useful to have run and weak to report. If you ran twelve PCR replicates per sample, the data to estimate a false positive rate is already sitting in your spreadsheet, unanalysed, and it costs nothing to fit.

Detect failure: a school club claiming to be Cochrane

What this review is not

This does not meet the definition a systematic review carries in medicine, and pretending otherwise would poison everything above.

A systematic review has a protocol registered before the search, a search strategy written out so somebody can repeat it exactly, two independent people screening every title and abstract with their disagreement rate reported, two independent people extracting every number, and a formal risk-of-bias assessment on each included study. We had one search, run iteratively as we learned which terms worked. No registered protocol. Extraction by whoever found the paper. We had no second screener. We computed no kappa. When two of us disagreed about whether Hinlo's fyke-net columns should be read across seasons, we argued about it at a table and wrote down the answer we agreed on, which is a fine way to run a club and is not a method.

What that means concretely, in decreasing order of how much it should worry you:

Our inclusion set is shaped by what we could get hold of. Every paper in the table was open access or otherwise readable without a subscription, except Schmelzle, which we included on secondary quotation and should perhaps not have. Extraction was single-pass, so any transcription error in the table survives, and the linked Python prints the source sentence for every row precisely so a reader can find one. The median-contrast rule was fixed in advance but the eligibility criteria were not fully fixed before we started looking, because we did not know until we had looked that per-unit detection probabilities for both methods would be this scarce. And we did no risk-of-bias assessment at all.

So what weight does 2.05 deserve? Less than a pooled estimate from a registered review, and more than zero. Treat it as a rough scale check: the filter is probably somewhere between a bit worse and several times better than the net, depending on the animal, and anybody who tells you the ratio with two significant figures is not reading the same literature. The number we would actually defend is \(\tau = 1.10\). The disagreement between studies is real and large. It measures the same whether or not our centre is off.

The false positive arithmetic in section 10 does not depend on our review: Bayes' rule and two rates. It holds exactly, no matter how badly we did the rest. If one thing here survives, let it be that.

References

  1. MacKenzie, D. I., Nichols, J. D., Lachman, G. B., Droege, S., Royle, J. A. & Langtimm, C. A. (2002). Estimating site occupancy rates when detection probabilities are less than one. Ecology 83, 2248–2255. doi:10.1890/0012-9658(2002)083[2248:ESORWD]2.0.CO;2
  2. Lahoz-Monfort, J. J., Guillera-Arroita, G. & Tingley, R. (2016). Statistical approaches to account for false-positive errors in environmental DNA samples. Molecular Ecology Resources 16, 673–685. doi:10.1111/1755-0998.12486
  3. Fediajevaite, J., Priestley, V., Arnold, R. & Savolainen, V. (2021). Meta-analysis shows that environmental DNA outperforms traditional surveys, but warrants better reporting standards. Ecology and Evolution 11, 4803–4815. doi:10.1002/ece3.7382
  4. DerSimonian, R. & Laird, N. (1986). Meta-analysis in clinical trials. Controlled Clinical Trials 7, 177–188. doi:10.1016/0197-2456(86)90046-2
  5. Higgins, J. P. T. & Thompson, S. G. (2002). Quantifying heterogeneity in a meta-analysis. Statistics in Medicine 21, 1539–1558. doi:10.1002/sim.1186
  6. Egger, M., Davey Smith, G., Schneider, M. & Minder, C. (1997). Bias in meta-analysis detected by a simple, graphical test. BMJ 315, 629–634. doi:10.1136/bmj.315.7109.629
  7. Valentini, A., Taberlet, P., Miaud, C., Civade, R., Herder, J., Thomsen, P. F., Bellemain, E., Besnard, A., Coissac, E., Boyer, F., Gaboriaud, C., Jean, P., Poulet, N., Roset, N., Copp, G. H., Geniez, P., Pont, D., Argillier, C., Baudoin, J.-M., Peroux, T., Crivelli, A. J., Olivier, A., Acqueberge, M., Le Brun, M., Møller, P. R., Willerslev, E. & Dejean, T. (2016). Next-generation monitoring of aquatic biodiversity using environmental DNA metabarcoding. Molecular Ecology 25, 929–942. doi:10.1111/mec.13428
  8. Biggs, J., Ewald, N., Valentini, A., Gaboriaud, C., Dejean, T., Griffiths, R. A., Foster, J., Wilkinson, J. W., Arnell, A., Brotherton, P., Williams, P. & Dunn, F. (2015). Using eDNA to develop a national citizen science-based monitoring programme for the great crested newt (Triturus cristatus). Biological Conservation 183, 19–28. doi:10.1016/j.biocon.2014.11.029
  9. Schmelzle, M. C. & Kinziger, A. P. (2016). Using occupancy modelling to compare environmental DNA to traditional field methods for regional-scale monitoring of an endangered aquatic species. Molecular Ecology Resources 16, 895–908. doi:10.1111/1755-0998.12501
  10. Hinlo, R., Furlan, E., Suitor, L. & Gleeson, D. (2017). Environmental DNA monitoring and management of invasive fish: comparison of eDNA and fyke netting. Management of Biological Invasions 8, 89–100. doi:10.3391/mbi.2017.8.1.09
  11. Eiler, A., Löfgren, A., Hjerne, O., Nordén, S. & Saetre, P. (2018). Environmental DNA (eDNA) detects the pool frog (Pelophylax lessonae) at times when traditional monitoring methods are insensitive. Scientific Reports 8, 5452. doi:10.1038/s41598-018-23740-5
  12. Rose, J. P., Wademan, C., Weir, S., Wood, J. S. & Todd, B. D. (2019). Traditional trapping methods outperform eDNA sampling for introduced semi-aquatic snakes. PLoS ONE 14, e0219244. doi:10.1371/journal.pone.0219244
  13. Akre, T. S., Parker, L. D., Ruther, E., Maldonado, J. E., Lemmon, L. & McInerney, N. R. (2019). Concurrent visual encounter sampling validates eDNA selectivity and sensitivity for the endangered wood turtle (Glyptemys insculpta). PLoS ONE 14, e0215586. doi:10.1371/journal.pone.0215586
  14. Moss, W. E., Harper, L. R., Davis, M. A., Goldberg, C. S., Smith, M. M. & Johnson, P. T. J. (2022). Navigating the trade-offs between environmental DNA and conventional field surveys for improved amphibian monitoring. Ecosphere 13, e3941. doi:10.1002/ecs2.3941
  15. Quilumbaquin, W., Carrera-González, A., Van der Heyden, C. & Ortega-Andrade, H. M. (2023). Environmental DNA and visual encounter surveys for amphibian biomonitoring in aquatic environments of the Ecuadorian Amazon. PeerJ 11, e15455. doi:10.7717/peerj.15455
  16. Li, W., Hou, X., Zhu, Y. et al. (2024). eDNA metabarcoding reveals the species–area relationship of amphibians on the Zhoushan Archipelago. Animals 14, 1519. doi:10.3390/ani14111519
  17. Di Girolamo, E. L., Jordan, M. A., Albers, G. & Bergeson, S. M. (2024). Comparing the effectiveness of environmental DNA and camera traps for surveying American mink (Neogale vison) in northeastern Indiana. PLoS ONE 19, e0310888. doi:10.1371/journal.pone.0310888
  18. Dougherty, M. M., MacDonald, A., York, G. & Post, D. M. (2025). Monitoring a keystone species (Alosa pseudoharengus) with environmental effects: a comparison with direct capture and environmental DNA. PLoS ONE 20, e0324385. doi:10.1371/journal.pone.0324385
  19. Smart, A. S., Tingley, R., Weeks, A. R., van Rooyen, A. R. & McCarthy, M. A. (2015). Environmental DNA sampling is more sensitive than a traditional survey technique for detecting an aquatic invader. Ecological Applications 25, 1944–1952. doi:10.1890/14-1751.1
  20. Hunter, M. E., Oyler-McCance, S. J., Dorazio, R. M., Fike, J. A., Smith, B. J., Hunter, C. T., Reed, R. N. & Hart, K. M. (2015). Environmental DNA (eDNA) sampling improves occurrence and detection estimates of invasive Burmese pythons. PLoS ONE 10, e0121655. doi:10.1371/journal.pone.0121655
  21. Lopes, C. M., Sasso, T., Valentini, A., Dejean, T., Martins, M., Zamudio, K. R. & Haddad, C. F. B. (2017). eDNA metabarcoding: a promising method for anuran surveys in highly diverse tropical forests. Molecular Ecology Resources 17, 904–914. doi:10.1111/1755-0998.12643