INTERACTIVE COMPANION · SYSTEMATIC REVIEW · CONSERVATION METHODS
The Detection Bench
Two benches sit below. The first re-pools the twelve studies live, so you can cut the set by taxon, by water volume, by what the paper counted as one survey, or by which model does the pooling, and watch the pooled diamond and the heterogeneity move together. The second bench is the one that matters. It has nothing to do with our review and everything to do with whether a positive result means anything.
On their default settings both reproduce the numbers in the article exactly.
Bench 1. Re-pooling the twelve
Each row is one study, one target species, and the log odds ratio between the per-unit detection probability of eDNA and that of the conventional comparator. Positive means the filter beat the net for one unit of effort. The diamond is the pooled estimate under a DerSimonian and Laird random effects model, and the pale bar under it is the prediction interval, which is where a thirteenth study's true effect would be expected to fall.
Watch two numbers as you filter. The pooled odds ratio moves less than you would expect. I squared barely moves at all, which is the whole finding.
Bench 2. What one positive is actually worth
Everything in bench 1 treats a positive result as the truth. This bench does not. Write p11 for the probability that an occupied site gives a positive and p10 for the probability that an unoccupied site gives one anyway, through contamination in the kit, contamination in the lab, a primer pair that cross-amplifies a related species, DNA carried down from a population upstream, or an index hop on a sequencing run. Then for a site with prior occupancy ψ, one positive sample gives
P(present | +) = ψ · p11 ψ · p11 + (1 − ψ) · p10
Which is Bayes' rule wearing a lab coat. Drag the occupancy slider left and watch the curve fall away underneath the marker. The false positive rate never changes while you do that. Nothing about the assay gets worse. The only thing that changes is how rare the animal is, and rarity is the exact condition under which people reach for this method.
- Pooled, all twelve
- odds ratio 2.05 (0.95 to 4.41)
- Heterogeneity
- I² = 80.5%, tau = 1.096 log odds
- Prediction interval
- odds ratio 0.15 to 27.4
- Default posterior
- 0.380 at ψ = 0.02, p10 = 0.02
- Two of two
- 0.948, at a sensitivity of 0.36
Three things to try
Set the taxon filter to reptiles and mammals. Three studies, pooled odds ratio 0.33, and I squared of exactly zero. Three papers that agree with each other perfectly while disagreeing with the amphibian papers by a factor of twelve. That is what real heterogeneity looks like when you catch it in the act, and no amount of careful extraction makes it go away.
Then leave the occupancy slider at 0.02 and push the false positive rate from 0.02 down to 0.005. The posterior climbs from 0.38 to 0.71 for a change of one and a half percentage points in a quantity nobody reports. Now do the opposite: leave p10 alone and improve p11 from 0.60 to 0.95, the best value in our whole table. The posterior goes from 0.38 to 0.49. Sensitivity is nearly worthless here and specificity is everything, which is the reverse of the intuition every survey protocol is built on.
Last, set r to 2 and then slide occupancy down to 0.005. The two-of-two rule holds up to about one site in a hundred and then starts failing too, because no fixed decision rule survives an arbitrarily rare target. At that point the only remaining moves are a better assay or a second line of evidence, and the second line of evidence is usually a net.