Science Journaling Club Founded 2024

EXPLAINER · EXPLAINER · SCIENTIFIC PRACTICE

How to Read a Scientific Paper Without Pretending You Understood It

Written jointly by the Science Journaling Club

Explainer · Peer-edited by the club review board · LaTeX source · Our calculation · Interactive model

Abstract Reading primary research is the one skill this club exists to teach. Almost nobody gets taught it. Here is the handout we wish somebody had pushed across the table on day one: what each part of a paper does, why experienced readers go to the figures before the prose, how to check whether the statistics carry the conclusion, and what peer review promises in writing. Honesty demanded that we measure the thing we were describing. With a text-analysis script we wrote ourselves, we scored 67 open-access research papers pulled from Europe PMC and all 298 sections of the club's own 27 published articles on one instrument. The median Results section runs 3,148 words against a median abstract of 206, a compression of 15.3 to one. Register changes harder than length does. Results sections favour reporting vocabulary over interpretation vocabulary 7.58 to one; abstracts flip it, with interpretation winning at 0.75 to one, a swing of about ten times. The popular charge that abstracts inflate their verbs did not survive our own test. Paired within each paper, abstract-minus-Results claim strength came out at +0.33 per thousand words, t = 0.16, n = 67, which is nothing at all. Abstracts do not shout. They change what kind of sentence they are writing, and they take the evidence away with them. Read for that.

Eleven at Night, Page Four

The first paper this club read together took five weeks.

Four pages. Five weeks. One hour a week, seven people, a projector that kept falling asleep, and at the end of it at least three of us could not have explained the second figure to a younger sibling. Nobody said so at the time. What everybody did instead was nod, in the specific way people nod when the alternative is admitting in front of a room that they are lost.

The nodding is the actual problem. Failing to understand a paper on first contact is ordinary. The people who wrote it feel it too, and go on feeling it for thirty years. Every scientist you have heard of has sat in front of a Methods section at eleven at night, read the same sentence about buffer concentrations four times running, gone to bed no wiser, and come back on a Tuesday morning to find it had quietly become obvious. The difference between them and a new club member is small. They stopped treating the experience as information about themselves.

A paper is a strange object, and it helps to know what kind. Call it a formal report. Written under severe length limits, aimed at perhaps two hundred people worldwide who already know the background, shaped by a century of conventions designed for archiving rather than for teaching, and cut down by reviewers working under length limits of their own. Every sentence carries load. Most carry it invisibly, until somebody shows you the joints. The prose has also, measurably, got harder: a study of more than 700,000 abstracts found readability declining steadily from 1881 to 2015, with the fraction of abstracts scoring below college reading level falling year on year [2]. You are not imagining the difficulty. It has been increasing for over a century.

So here is what this article is. The handout we should have written before that first meeting, covering the anatomy of a paper, the order to read it in, how to test an abstract against its own results, what the statistics have to do before a conclusion is earned, what peer review actually certifies, how to recognise a paper that is mostly a press release, and what to do at the moment you hit a paragraph that will not open. Four of its sections are a measurement. A guide to reading evidence that then served you opinion would be a poor joke.

Six Rooms, and What Each One Is For

The shape of a modern paper has a name and a birthday. The name is IMRaD, for Introduction, Methods, Results and Discussion. It went from one option among several to near-universal in medical journals between roughly 1940 and 1980; a survey of four leading journals found IMRaD used in none of the papers sampled in 1935 and in essentially all of them by 1985 [1]. Before that, papers were written more like letters, or like essays, or like whatever the author felt like that month. The standardisation was a filing decision as much as a scientific one. Every paper you pick up now has its parts in the same rooms.

Knowing what each room is for is most of the skill. Each is written under different rules, and each earns its own particular kind of suspicion.

The title is a search string, written so the right specialist finds the paper and, increasingly, so the paper gets clicked at all. State a conclusion up there and you have promised something. The rest of the paper pays.

The abstract summarises the paper for people who will never read it. Nearly everybody, then. Ours above runs 210 words, which is typical. Across our 67 papers the median abstract ran 206 words and summarised a median Results section of 3,148, which makes it the most-read and least-checkable text in science, and the only part of the paper most readers will ever carry in their heads.

The introduction tells you what problem exists and why this work is the next move. An argument dressed as background. The citations were picked to build it, so read the section for the question and read it knowing the literature is being curated in front of you.

The methods section gets skipped by everybody. It also decides whether anything else in the paper is true. Written for replication rather than comprehension, it reads like assembly instructions for furniture you do not own, which is roughly the effect intended, since the reader it imagines is a rival trying to repeat the work rather than a student trying to follow it. The sample size lives here. So does the thing actually measured, often not the thing the title talks about.

The results section reports what happened. Past tense, ideally without a word about what any of it means. The longest section in a typical paper, and the one that carries the figures. Median 3,148 words in our sample, maximum 12,675.

The discussion lets the authors interpret, speculate, and say why this matters. Most readable part of most papers. Also the least binding. Treat it as the authors' best case rather than a finding.

Two more rooms hide behind the others. The limitations paragraph, usually near the end of the Discussion, is where honest authors tell you exactly how to attack their own work, which is why we read it first on any evening when the reading has to be quick. The supplementary material is where awkward data goes to be technically available.

Here is that anatomy with our own measurements attached. Every number in the middle four columns comes from the 67 papers described in §4, scored by the script linked at the top of this page.

Section Median words Hedges per 1,000 Report : interpret Written to Read it for
Abstract 206 5.16 0.75 persuade a stranger in ninety seconds the claim, so you can go and test it
Introduction 704 5.58 1.07 establish that the question is open the question, and who is being left out
Methods 1,652 2.62 9.00 let a rival repeat the work sample size, and what was actually measured
Results 3,148 4.16 7.58 report outcomes without arguing effect sizes, intervals, figures, and outcomes measured but not reported
Discussion 1,194 11.30 1.27 interpret, and place the work the limitations paragraph
Club article section 358 3.24 2.00 explain to somebody outside the field whether we cited the paper or a summary of it

Read the last two numeric columns as a pair. Hedging density counts words such as may, might, could and appears per thousand words, and the report-to-interpret ratio compares the vocabulary of saying what happened against the vocabulary of saying what it means. Methods and Results are overwhelmingly reporting text with little hedging. The Discussion hedges at 11.30 per thousand, nearly three times the Results rate, which is exactly right: speculation should announce itself. The abstract is the interesting row. §5 is about why.

Read the Figures Before You Read the Words

Try this on the next paper you open, even if it feels backwards.

Skip the abstract. Go to the first figure. Read its caption twice, slowly, because a caption in a good paper is a complete miniature argument and in a bad paper the caption is where the argument quietly stops being supported. Look at the axes. Find out what is being plotted against what. Check the units. Check whether zero sits on the scale. Count the data points if you can. Then decide what you think the figure shows. Do it before you have read a word of the authors' prose.

Now read the Results text about that figure. Do the authors agree with you?

That disagreement, when it happens, is the single most useful thing an inexperienced reader can find, and it is only available to you if you form your own impression first. Read the prose first and you will see what you were told to see. Scientists' honesty has nothing to do with it. The effect is about how reading works.

The order matters for a second reason, which is triage. Most papers you open are not worth finishing. A working scientist decides in about four minutes, and they do it by checking the figures and the sample size, because those two things determine whether the conclusions can be true regardless of how well the Discussion is written [3]. Figure 1 lays the two orders side by side: the front-to-back march that a new reader attempts, and the loop an experienced reader runs.

ORDER A · FRONT TO BACK ORDER B · WHAT READERS ACTUALLY DO Abstract 206 w Introduction 704 w Methods 1652 w Results + figures live here 3148 w Discussion 1194 w limitations buried here block heights are proportional to median length, n = 67 1 2 3 4 5 Title + figure captions what is plotted against what? Methods, two facts only sample size; what was measured Results, around the figures do they say what you saw? Limitations paragraph the authors' own attack list Introduction, if still reading why anyone asked this Abstract, last now you can grade it 1 2 3 4 5 6 stop here, most of the time triage is the normal outcome
Figure 1. Two ways through the same object. On the left, the paper as printed, with block heights proportional to the median section lengths we measured across 67 open-access papers. Order A reads it as a book. Order B is the loop most working readers run: captions and axes first, two facts from the Methods, the Results text only where it touches a figure, then the limitations, and the abstract last of all, once you can actually grade it. The dashed ramp is the part nobody mentions. Most papers get abandoned after step two, and abandoning them is the correct decision rather than a failure of nerve.
Query from the review board

Is reading the abstract last dishonest advice? The abstract is the only part most people will ever read.

Answer: we recommend it for papers you have decided to read properly, perhaps one in twenty. For the other nineteen the abstract is all you get. §5 is about how much that costs you.

Plain Arithmetic

This section states numbers and their sources. No argument lives in it.

Two bodies of text were measured. Corpus A is the club's own output: every article on this site except the template and this page, which comes to 27 articles, 298 sections marked by an h2 heading, and 115,159 words of body prose. Figure captions, tables, margin notes and reference lists were excluded. Corpus B is 67 research papers, retained from 96 candidates drawn from the Europe PMC open-access subset. The frame was CC-BY research articles published between 1 January 2022 and 31 December 2024 in eight journals, twelve most-cited articles per journal; the 96 identifiers sit in the script, so the selection can be repeated or disputed. A candidate was kept only if the full-text XML held both an abstract and a top-level section whose title began with "Result". Twenty-nine candidates failed that test. Nothing was replaced. Corpus B supplied 231,448 words of Results text and 15,907 of abstract.

Figure 2 draws the section lengths to scale. Median section lengths in Corpus B, in words: abstract 206, introduction 704, methods 1,652, results 3,148, discussion 1,194. The ratio of median Results to median abstract is 15.3. The tenth and ninetieth percentiles for the abstract are 147 and 354 words; for the Results, 967 and 5,824.

Mean sentence length: 22.9 words in abstracts, 24.6 in Results, 24.9 in Methods, 26.9 in Discussions. Corpus A averages 16.1. Standard deviations are 10.8 for abstracts and 18.5 for Results, on 694 and 9,391 sentences respectively.

Hedging density, in hits per thousand words: abstract 5.16, introduction 5.58, methods 2.62, results 4.16, discussion 11.30, Corpus A 3.24. The lexicon holds 39 forms. The ten commonest hedges across abstracts, Results and Discussions were may at 401 occurrences, could at 285, potential at 162, likely at 134, possible at 124, might at 119, suggest at 89, suggests at 57, potentially at 55 and relatively at 53.

Two checks were run against values fixed in advance. The first compares the tokeniser against a hand count of a five-sentence passage: hand 5 sentences, 30 word tokens and 3 hedges, measured 5, 30 and 3. The second compares a bootstrap standard error against the analytic value \(s/\sqrt{n}\) from the central limit theorem. On the 9,391 Results sentences, with standard deviation 18.4525 words and a resample size of 200, the analytic standard error is 1.30479 and the bootstrap estimate is 1.30659, a ratio of 1.0014. Rescaled to the full sample the two read 0.19041 and 0.19068. Both checks print measured beside expected in the output file.

abstract 206 introduction 704 methods 1,652 results 3,148 discussion 1,194 club section 358 0 1k 2k 3k 4k 5k 6k words (bars = median, whiskers = p10 to p90) 15.3× compression Corpus B: 67 papers, Europe PMC open access, 2022–2024 · Corpus A: 298 sections, 27 club articles
Figure 2. Median section length, 67 papers. The abstract bar is 15 pixels wide at this scale and the Results bar is 230. Everything the abstract claims rests on the long bar, and the reader who stops at the short one has seen 6.1% of the evidence-bearing text and none of the numbers underneath it. The club's own sections, median 358 words, are shorter than a typical Introduction, which is a choice about audience rather than a virtue.

Claim-strength scores per thousand words, on a graded lexicon of 81 forms weighted from -2 to +2: abstracts mean 10.33 and median 11.15; Results mean 10.00 and median 9.57; Discussions mean -3.49 and median -6.72. The paired difference within each paper, abstract minus Results, has mean +0.33, standard deviation 17.32, standard error 2.12, and t = 0.16 on 66 degrees of freedom. In 35 of 67 papers the abstract scored higher; in 32 it scored lower. Corpus A scored 1.13 for abstracts and 1.38 for body prose, a gap of −0.25.

All printed output is in reading-a-paper-output.txt. The script is reading-a-paper.py.

The Register Switch, and the Overclaim That Was Not There

We went in expecting to catch abstracts red-handed.

The claim is everywhere, including in places we trust: abstracts inflate, they promise more than the data delivers, they are the marketing layer of a scientific paper. Real evidence stands behind it. Human raters working through randomised trials with negative primary outcomes found spin, meaning presentation that distracts from a non-significant result, in the conclusions of 58% of abstracts they examined [4], and a systematic review of spin research found it reported across a wide sweep of the biomedical literature [5]. The same group has catalogued how distortion enters at each stage of the chain from result to reader, and how little of it is deliberate [6]. Word choice has been drifting too: the frequency of positive words such as novel, unprecedented and innovative in PubMed abstracts rose roughly ninefold between 1974 and 2014 [7].

So we built a graded claim-strength lexicon, scored every abstract and every Results section, paired them within each paper so that topic, authorship, field and house style cancel out, and ran the difference.

The difference came out at +0.33 per thousand words. The standard error was 2.12. Divide one by the other and you get a t of 0.16, meaning the measured gap is about one fifteenth the size of its own uncertainty, and meaning that a world with no difference in it at all would hand you a gap this big or bigger the overwhelming majority of the time. Thirty-five of 67 abstracts scored more assertive than their own Results, and 32 scored less. Call it 52% against 48%. A coin.

interpretation wins reporting wins abstract 0.75 : 1 introduction methods 9.00 : 1 results 7.58 : 1 discussion 1.27 : 1, hedges hardest club prose same side of the line 0.5 1 2 5 10 reporting words per interpretation word (log scale) 0 2 4 6 8 10 12 hedges per 1,000 words 67 papers, Europe PMC open access · club prose = 298 sections of our own WHERE EACH SECTION LIVES
Figure 3. From the club calculation. Each point is one section type, placed by how much it reports against how much it interprets on the horizontal axis, and by how hard it hedges on the vertical. Methods and Results sit far to the right: they are reporting text. The Discussion sits high on the left, interpreting and hedging at 11.30 per thousand words, which is the honest signature of a section that knows it is speculating. The abstract sits low on the left. It is on the Discussion's side of the line, writing interpretation, while hedging at only 5.16 and carrying the authority of a result. That position, rather than any inflated verb, is what makes abstracts misleading.

Look at where the abstract landed. On the reporting-to-interpretation measure it scores 0.75, using more interpretation words than reporting words, which puts it on the Discussion's side of the line at 1.27 and nowhere near the Results at 7.58. Roughly a factor of ten separates the section a reader reaches first from the section holding the evidence. And it hedges at 5.16 per thousand against the Discussion's 11.30. The Discussion's job, then, at less than half the Discussion's caution.

The finding we did get is more useful. An abstract is a compressed Discussion wearing the Results' uniform. Nobody needs to exaggerate a verb for that to mislead you, because the misleading happens a level up, in the question of what kind of sentence is being written at all.

0 +40 +20 −20 −40 mean 10.33 mean 10.00 ABSTRACT RESULTS claim strength per 1,000 words, one thread per paper, n = 67 paired difference +0.33 ± 2.12 t = 0.16, 66 df 35 up, 32 down one paper at +53.06 one paper at −37.50
Figure 4. The null we did not want. Each faint thread joins one paper's abstract to its own Results section, scored on the same 81-form claim lexicon. The threads cross both ways and the two means, 10.33 and 10.00, are almost on top of each other. Individual papers swing hugely, from +46.06 to -37.06, which is why the standard deviation of the paired difference is 17.32 and the mean is statistically invisible. Read this as a warning against using a population tendency to convict a particular paper, and read Figure 3 for the effect that is actually there.

Can the Statistics Hold the Conclusion Up?

Six questions, in the order we ask them. None requires a statistics course.

  1. What is n, and n of what? Find the sample size in the Methods. Then check what it counts. A study reporting 240 measurements from eight animals has an n of eight, for any claim about animals. That swap is the commonest quiet error in print. One sentence usually gives it away.
  2. What is the effect size, in units you can picture? A p-value, meaning the probability of seeing data at least this extreme if the effect were truly absent, tells you nothing about how big anything is. "Significantly reduced" is compatible with a reduction of 0.3%. Look for the number and its units, and if the paper does not give you one, that is itself the finding.
  3. Is there a confidence interval, and what does its far end allow? A 95% confidence interval spans the effect sizes the data cannot rule out. If it runs from "a 2% improvement" to "a 60% improvement", the paper has not established that the effect is large; it has established that it is probably positive. The far end is where the honesty lives.
  4. Was the study big enough to see what it claims to have seen? Statistical power is the chance of detecting an effect that is genuinely there. The median power across the neuroscience literature has been estimated at between 8% and 31% [13], which means most studies in that field were structurally unable to find what they were looking for, and a small study that reports a large effect is more likely to be reporting noise than a breakthrough.
  5. How many things were tested? Run twenty independent tests at the usual 5% threshold and one will come up apparently positive by chance alone, which is why a paper that measured forty outcomes and reports the three that worked is a lottery result rather than a discovery. A pre-registered analysis plan, written before the data were collected, carries more weight for exactly this reason.
  6. Do the numbers in the text agree with each other? They often do not. An automated check of over 250,000 p-values in eight psychology journals found that half of the papers contained at least one p-value inconsistent with its own reported test statistic and degrees of freedom, and about one in eight contained an inconsistency large enough to change a conclusion [14].

Two warnings about the p-value. Of every number in a paper, this is the one most likely to be waved at you. A p-value is not the probability that the hypothesis is true. Nor is it the probability that the result was a fluke. After a century of that misreading, the American Statistical Association took the unusual step of publishing a formal statement saying so [10]. One widely cited catalogue lists twenty-five distinct ways that p-values, intervals and power get misinterpreted, including by the people who teach them for a living [11]. So the confusion is not a verdict on your own slowness. In 2019 more than eight hundred scientists signed a comment in Nature arguing that the whole category of "statistically significant" should be retired, on the grounds that dichotomising a continuous measure of evidence at 0.05 causes more errors than it prevents [12].

You do not need to settle that argument. You need to notice when a paper's conclusion rests entirely on one side of it.

Query from the review board

Should a new member really be checking p-values against test statistics by hand?

No. Ask questions 1, 2 and 5 every time, because they need arithmetic rather than training. Question 6 is there so you know that the error rate is real and that finding one does not mean you have misread something.

What Peer Review Promises, in Writing

Peer review is the part of science most trusted by people outside it and most argued about by people inside it. Precision about the mechanism helps.

An editor sends a submitted manuscript to two or three researchers in a related field, who read it in their own time, usually unpaid and usually anonymously, and send back a recommendation with comments. The authors revise. The editor decides. No raw data changes hands. Everything anybody believes about peer review has to fit inside that description.

What it reliably does: it filters out work that is obviously unsound to a specialist, and it forces clarity by making authors answer awkward questions before publication. Those are real goods and the alternative is worse.

Verification is the part it does not do. Reviewers almost never see the raw data, and rerunning the analysis is rarer still. Nobody repeats the experiment. A Cochrane review of the evidence on editorial peer review concluded that despite its central place in the system, there was little empirical evidence to support its use as a mechanism for ensuring quality [17]. When researchers deliberately inserted errors into short papers and sent them to reviewers, the reviewers found a median of between two and three of the nine major errors planted, and training the reviewers produced only a small and short-lived improvement [18]. Reviewers are colleagues reading quickly, which is a different job from auditing.

Two consequences follow for you as a reader, and they pull in opposite directions.

First: "peer-reviewed" marks a floor, not a certificate. Large replication efforts have repeatedly found that a substantial fraction of published, peer-reviewed findings do not hold up when the study is run again: one coordinated attempt to repeat 100 psychology studies reproduced the original result in roughly a third to a half of cases depending on the criterion used [16]. Much of that is ordinary statistical fragility rather than misconduct, and it is the predicted consequence of a system where small studies, selective reporting, flexible analysis and publication bias interact [15].

Second: the absence of peer review disqualifies nothing either. Preprints, which are manuscripts posted publicly before review, are how a great deal of physics and increasingly of biology now circulates, and some of them are excellent. What changes is where the burden sits. With a peer-reviewed paper you are trusting that two specialists saw no obvious problem. With a preprint you are the specialist. Neither state relieves you of reading the Methods.

The Paper That Is Mostly a Press Release

Some papers are built to be covered rather than read. Fraud is rarely the mechanism. They are optimised, in the way a headline is optimised, and the optimisation happens at every step of a chain that has been traced end to end.

The chain runs like this: a result, a press release written by a university communications office, a news story written from the release, a social post written from the news story. A study of 462 press releases and the papers and news stories attached to them found that exaggeration in the news was overwhelmingly predicted by exaggeration already present in the press release, with 40% of releases containing exaggerated advice, 33% containing exaggerated causal claims from correlational data, and 36% exaggerating the inference from animals to humans [8]. When the release was accurate, the news was mostly accurate. A separate analysis of randomised trials found spin in 47% of press releases and 51% of associated news items, and that spin in the abstract predicted spin downstream [9].

So the tell is in the paper.

Specimen · abstract sentence

Our findings demonstrate that compound X prevents cognitive decline, opening a new therapeutic avenue for Alzheimer's disease.

Three marks. "Demonstrate" in an abstract usually covers a single experiment. "Prevents" is a causal verb doing work the design cannot support. "Opening a new therapeutic avenue" appears in no Results section ever written, because it is not a result.

Specimen · the same finding, Results section

Mice receiving compound X showed a 14% smaller decrement in maze completion time at 12 weeks than vehicle-treated controls (n = 9 per group, p = 0.04).

Nine mice per group. A 14% difference on one behavioural measure at one time point. Nothing in this sentence is dishonest, and nothing in it supports the word "prevents".

Both of those are constructed, but the gap between them is the ordinary gap, and once you have seen it you cannot stop seeing it. The specific things we check, in the order they are quickest to check:

  1. Does the abstract's conclusion name a species, a population or a condition that the Methods never mention? Mice becoming people is the commonest single jump.
  2. Does a causal verb appear over an observational design? Words such as causes, prevents, drives and protects need an intervention somewhere in the Methods, and if the Methods describe a cohort followed over time with nothing done to it, the verb has outrun the design.
  3. Is the headline number a relative change with no absolute change beside it? "Doubles the risk" covers two cases per hundred thousand becoming four.
  4. Does the paper's own press release exist, and does it say something the paper does not? University releases sit one click away. Comparing them is the fastest ten seconds in science journalism.
  5. Is the most quotable sentence in the paper sitting in the Discussion's final paragraph? That paragraph gets written to be quoted. Nothing else in the document is bound so loosely to evidence.

The Paragraph You Cannot Follow

It happens on page three, usually, and it happens to everybody.

You have been reading along fine, the argument has been holding, and then five sentences arrive that might as well be in another alphabet. Something about a generalised linear mixed model with crossed random effects, then something about an orthogonal projection onto a residual subspace, and every word between them one you know, arranged in an order you do not. The temptation at this moment is enormous. Keep the eyes moving. Hope the next paragraph lets you back in.

Here is what we do instead, and it is not heroic.

  1. Name the obstacle out loud. Write it in the margin. "I do not know what a random effect is." Four seconds of work. A vague sense of failure becomes one specific lookup. Most of the time the paragraph turns out to be hard for exactly one reason.
  2. Decide whether it is load-bearing. Ask what the paper would lose if the paragraph were deleted. If it justifies a choice of statistical model and the conclusion does not turn on that choice, note it and move on. If it is the step where the data become the claim, you have to stay.
  3. Read the sentence before and the sentence after. Authors very often state the point of a technical passage in plain language on either side of it, which is a habit of writing rather than a kindness, since the plain sentence is the one a reviewer will check. The difficult middle is the proof. The easy edges are the claim.
  4. Find the same idea somewhere it is explained rather than used. A review article, a textbook chapter, a good encyclopedia entry, somebody's lecture notes. Papers use concepts; they do not teach them, and expecting a paper to teach you is like expecting a contract to teach you law.
  5. Write the question down and bring it to a person. The club meets for exactly this. Half of what any of us knows came from somebody else saying "oh, that just means they measured each animal more than once."
  6. Allow yourself to leave it unresolved and say so. In our written notes we mark unresolved passages with a bracketed question rather than smoothing over them, and those brackets are the most useful thing in the file when we come back.

The rule underneath all six steps: an admitted gap is a working state, and a concealed gap is a broken one. The moment you decide to pretend, you have stopped being able to learn from the paper, because you can no longer tell which parts you actually have.

One more encouragement, empirical rather than pastoral. The difficulty is not all in you. Scientific prose has been getting measurably less readable for more than a century, with the trend visible across 12 million abstracts and driven partly by a rising density of specialist vocabulary and general scientific jargon [2]. Somebody reading a 1950 paper in the same field would have an easier time than you are having. The object changed under you. You did not get worse.

Working Notes From the Club Table

Session 1 · Tuesday · the lexicon fight

Question on the whiteboard: how do you measure "overclaiming" without just asking somebody's opinion? Answer proposed: count words. Objection raised immediately, and correctly, that counting words cannot read a sentence. "We did not show that X causes Y" scores as assertive under any list-based method, because the list sees show and causes and never sees the not. Nobody solved this. We wrote it into the limitations and kept going, on the grounds that a crude instrument applied identically to both sides of a paired comparison is still informative about the difference.

Longer argument about whether significant should count as an assertive word. A technical term with a precise meaning, and the most abused word in the language of science. We gave it +1, half the weight of demonstrate, and recorded that as a judgement call rather than a fact.

Session 2 · Thursday · the validation we failed

We wrote a five-sentence test passage and three of us counted it by hand before running the code, which is the right order. Hand count: 5 sentences, 38 words, 3 hedges. Code: 5 sentences, 30 words, 3 hedges. Eight words missing and a long silence.

The code was right. Two of us had counted "12" and "4.2" as words when the tokeniser's stated rule is letter-runs only, and none of us had noticed that "i.e." contributes two one-letter tokens. The hand count held three separate errors and they partly cancelled. That single morning is the entire argument for validating against something fixed in advance, and the wrong number still sits in the script's comments where anybody can find it.

Session 3 · Sunday · when the result refused

Ran the paired comparison expecting a clear positive gap. Got +0.33. The standard error was 2.12. Somebody asked whether we should try a different lexicon. Somebody else pointed out that we had only decided the lexicon was correct when we thought it was going to give us the answer we wanted. That ended the discussion faster than any argument about statistics could have.

Then the better idea arrived, from the person who had been quietly staring at the ratio column. The abstract's report-to-interpret ratio was 0.75 and the Results' was 7.58, sitting in the same table we had been reading past for two weeks. The effect was there. Sitting in plain view. In a different variable all along.

Session 4 · the following Tuesday · scoring ourselves

Ran the whole instrument over our own 27 articles so that we would be in the table too. Club prose hedges at 3.24 per thousand against the papers' 4.16 in Results, our sentences average 16.1 words against their 24.6, and our report-to-interpret ratio is 2.00, which puts us squarely between a Results section and a Discussion. Roughly what an explainer should be, and we are not going to pretend we planned it. One number did sting a bit. Our own abstracts score 1.13 on the claim lexicon against 1.38 for our body prose, so on our own instrument we are marginally more cautious out front than we are inside, which is the opposite of the effect we set out to catch other people committing.

What People Get Wrong, and Why the Wrong Version Is Nicer

Four misconceptions, all of them comfortable.

The first is that scientists hype their abstracts. We believed this version going in, and our own measurement declined to support it: paired within paper, claim-verb strength barely moves between abstract and Results. The appeal of the hype story is that it makes the problem intentional, and therefore fixable by suspicion, which is a cheap repair and a flattering one. Be cynical enough and you are safe. What we found instead is duller and harder to defend against, because the abstract misleads by being a different genre rather than by lying, and no amount of general distrust will tell you which 3,148 words of Results have gone missing behind those 206.

The second is that peer review means it is true. Two or three people read it quickly and saw nothing obviously wrong. The whole claim, and worth something. The wrong version is appealing because it lets a non-specialist outsource the judgement entirely to somebody qualified, which would be a wonderful arrangement if it worked the way people imagine it working. The wrong version appeals in the opposite direction too, to people who want to dismiss a finding, because "peer review is broken" is a much easier sentence to say than "here is what is wrong with the third figure."

The third is that difficulty is a personal verdict. If a paper is hard, the verdict lands on you: not clever enough, or short of the prerequisites, and either way not somebody who belongs in the room. This one is nastier than the others because it feels like humility. Bleak comfort: a complete explanation that requires nothing further from you, and permission to stop. The measured truth is that papers are written in a compressed dialect for an audience of roughly two hundred people, and that the dialect has been getting harder for a hundred and thirty years [2]. A paper being difficult is information about the paper.

The fourth is that reading a paper means reading all of it, in order. Beginning at the abstract and marching to the references feels like respect. It feels like thoroughness. Neither, as it turns out. That order guarantees you meet the authors' interpretation before you have seen a single data point, and it explains why people finish a paper convinced of a conclusion whose supporting figure they never really looked at. The appeal is straightforward. Front-to-back is what reading has meant since you were five. Abandoning it feels like cheating, and it is the single highest-return habit change available to a new reader.

One thread runs through all four. Each wrong version converts an ongoing obligation into a single settled fact, and settled facts are restful. Suspicion replaces checking. Certification replaces reading. A verdict about yourself replaces a question about the text.

The Strongest Case Against This Article

Now the other side, put as well as we can put it, because an explainer that never argues with itself is a sermon.

Our null result is probably wrong, and the literature is probably right. The serious objection, and it cuts at our own headline. Human raters, working slowly, with training, reading whole abstracts in context, find spin in roughly half of the abstracts they assess [4, 5, 9]. We found nothing, using a word list, on 67 papers, in a frame skewed towards methods and tools in the life sciences, with an instrument that cannot detect negation and cannot tell a hedge placed honestly from a hedge placed to cover a weak result. When a careful human method and a crude automated method disagree, the automated method is the one that should lose, and ours is unmistakably the crude one. The correct reading of our Figure 4 is that our instrument failed to detect an effect that better instruments do detect, and we would be overclaiming, in exactly the way this article warns about, if we said otherwise. What survives is Figure 3, because a register difference of that size is too large for lexicon noise to manufacture.

Figures-first reading manufactures confident misreadings. A figure separated from its Methods is close to an inkblot. You can look at a beautiful scatter plot and form a firm impression that survives contact with the prose, and be wrong because the axis is log-scaled, or because the points are technical replicates rather than independent samples, or because a control condition that would flatten the whole effect is described two pages away. Reading order B in Figure 1 puts a novice in front of the most persuasive object in the paper at the moment they have the least ability to resist it. The defence is thin but real: order B sends you to the Methods for sample size at step two, before the Results text, and that is precisely to stop this. If you drop that step the criticism lands completely.

The advice is aimed at the wrong failure mode. A press-release paper is engineered to survive figure-first reading, because the figures are the part that gets professionally designed, sometimes by somebody whose whole job is making a modest result look inevitable. The reader most likely to be fooled by a slick paper is the reader who has learned to trust figures and has not yet learned that a well-made figure of a badly-designed experiment is still a well-made figure.

And the checklists may do harm. Picture the reader who learns six statistical questions, applies them mechanically, finds a small n, and dismisses a careful qualitative study that was never trying to be a trial. Checklists produce false confidence in their users at least as easily as abstracts produce false confidence in theirs, and a reader armed with six questions and no judgement is a new hazard rather than a solved problem. We have no clean answer to this. The only honest mitigation is the one in §9: the questions are for deciding what you still need to find out, and a checklist that terminates your thinking has been used backwards.

What survives all four objections is small. We think it is right. Read the figures with your own eyes, before anybody's interpretation of them. Count what was measured. Then count what got reported. Know which room you are standing in, because the rules differ in each. And when the paragraph will not open, say so, in writing, to somebody else.

The whole method, there. It took this club five weeks to read four pages, and the five weeks were not wasted, because what we were actually building was the habit of admitting where we were.

References

  1. Sollaci, L. B. & Pereira, M. G. (2004). The introduction, methods, results, and discussion (IMRAD) structure: a fifty-year survey. Journal of the Medical Library Association 92(3), 364–367. PMCID:PMC442179
  2. Plavén-Sigray, P., Matheson, G. J., Schiffler, B. C. & Thompson, W. H. (2017). The readability of scientific texts is decreasing over time. eLife 6, e27725. doi:10.7554/eLife.27725
  3. Pain, E. (2016). How to (seriously) read a scientific paper. Science Careers. doi:10.1126/science.caredit.a1600047
  4. Boutron, I., Dutton, S., Ravaud, P. & Altman, D. G. (2010). Reporting and interpretation of randomized controlled trials with statistically nonsignificant results for primary outcomes. JAMA 303(20), 2058–2064. doi:10.1001/jama.2010.651
  5. Chiu, K., Grundy, Q. & Bero, L. (2017). "Spin" in published biomedical literature: a methodological systematic review. PLOS Biology 15(9), e2002173. doi:10.1371/journal.pbio.2002173
  6. Boutron, I. & Ravaud, P. (2018). Misrepresentation and distortion of research in biomedical literature. Proceedings of the National Academy of Sciences 115(11), 2613–2619. doi:10.1073/pnas.1710755115
  7. Vinkers, C. H., Tijdink, J. K. & Otte, W. M. (2015). Use of positive and negative words in scientific PubMed abstracts between 1974 and 2014: retrospective analysis. BMJ 351, h6467. doi:10.1136/bmj.h6467
  8. Sumner, P., Vivian-Griffiths, S., Boivin, J., Williams, A., Venetis, C. A., Davies, A., Ogden, J., Whelan, L., Hughes, B., Dalton, B., Boy, F. & Chambers, C. D. (2014). The association between exaggeration in health related science news and academic press releases: retrospective observational study. BMJ 349, g7015. doi:10.1136/bmj.g7015
  9. Yavchitz, A., Boutron, I., Bafeta, A., Marroun, I., Charles, P., Mantz, J. & Ravaud, P. (2012). Misrepresentation of randomized controlled trials in press releases and news coverage: a cohort study. PLoS Medicine 9(9), e1001308. doi:10.1371/journal.pmed.1001308
  10. Wasserstein, R. L. & Lazar, N. A. (2016). The ASA statement on p-values: context, process, and purpose. The American Statistician 70(2), 129–133. doi:10.1080/00031305.2016.1154108
  11. Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N. & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations. European Journal of Epidemiology 31, 337–350. doi:10.1007/s10654-016-0149-3
  12. Amrhein, V., Greenland, S. & McShane, B. (2019). Scientists rise up against statistical significance. Nature 567, 305–307. doi:10.1038/d41586-019-00857-9
  13. Button, K. S., Ioannidis, J. P. A., Mokrysz, C., Nosek, B. A., Flint, J., Robinson, E. S. J. & Munafò, M. R. (2013). Power failure: why small sample size undermines the reliability of neuroscience. Nature Reviews Neuroscience 14, 365–376. doi:10.1038/nrn3475
  14. Nuijten, M. B., Hartgerink, C. H. J., van Assen, M. A. L. M., Epskamp, S. & Wicherts, J. M. (2016). The prevalence of statistical reporting errors in psychology (1985–2013). Behavior Research Methods 48, 1205–1226. doi:10.3758/s13428-015-0664-2
  15. Ioannidis, J. P. A. (2005). Why most published research findings are false. PLoS Medicine 2(8), e124. doi:10.1371/journal.pmed.0020124
  16. Open Science Collaboration (2015). Estimating the reproducibility of psychological science. Science 349, aac4716. doi:10.1126/science.aac4716
  17. Jefferson, T., Rudin, M., Brodney Folse, S. & Davidoff, F. (2007). Editorial peer review for improving the quality of reports of biomedical studies. Cochrane Database of Systematic Reviews 2, MR000016. doi:10.1002/14651858.MR000016.pub3
  18. Schroter, S., Black, N., Evans, S., Godlee, F., Osorio, L. & Smith, R. (2008). What errors do peer reviewers detect, and does training improve their ability to detect them? Journal of the Royal Society of Medicine 101(10), 507–514. doi:10.1258/jrsm.2008.080062