A scientific paper can be careful, useful and wrong. It can also be correct about a narrow question and misleading when translated into a headline, a policy or advice for someone outside the study. Reading critically does not mean searching for a fatal flaw. It means discovering exactly what the paper can support.
The abstract is designed to compress. Compression hides choices: who was included, what counted as an outcome, when it was measured, which comparison was made and how much uncertainty surrounds the estimate. Begin with a map of those choices before accepting the conclusion.
First decide how deeply to read
Not every paper deserves an hour. Use three levels. For a low-stakes curiosity, read the question, design, main estimate and limitations. For a decision affecting money, health, work or policy, read methods, outcome definitions, absolute effects and conflicts. For a claim you will publish, teach or use to guide other people, inspect the protocol or registration when available, supplementary material, cited prior evidence and relevant systematic reviews.
This guide is a literacy protocol, not a substitute for domain expertise or clinical advice. Specialized methods can contain assumptions that a general reader will not recognize. The honest endpoint may be “I understand the claim but cannot independently evaluate this analysis.”

1. Reconstruct the question
Write one sentence using five elements: population, exposure or intervention, comparator, outcome and time. “Does the program work?” becomes “Among first-year students at these universities, did access to this eight-week program, compared with usual services, change a defined anxiety score at twelve weeks?”
If you cannot fill a field, mark it unknown rather than guessing. Notice surrogate outcomes. A change in a laboratory marker is not automatically a change in symptoms, survival or quality of life. Notice composite outcomes. Combining several events can conceal that the result is driven by the least important or most frequent component.
2. Identify what the design can distinguish
Randomized trials can reduce confounding when allocation, follow-up and analysis are handled well. Observational studies can reveal associations in large or realistic populations, but alternative explanations may remain. Cross-sectional studies measure a moment and usually cannot establish which condition came first. Case-control studies look backward from outcomes and depend on suitable controls and measurement. Qualitative work can illuminate experience and mechanism without estimating population effect sizes.
Do not rank designs with one universal ladder. Match design to question. Randomizing a suspected harm may be unethical. A randomized trial may be too short or selective to reveal rare harms. A strong observational design can answer a question that a weak trial cannot. Read the methods to learn what comparison was actually created.
3. Inspect selection, exclusion and loss
Who could enter? Who was excluded? Who agreed? Who completed follow-up? A study of healthy volunteers, users of one platform or patients at specialist centers may not transport cleanly to older adults, children, people with multiple conditions or communities with different access.
Attrition matters when departure is related to outcome or treatment. Ask whether groups lost similar proportions, why participants left and how missing data were handled. A final sample can look balanced after the people harmed, burdened or unconvinced have disappeared.
4. Read how variables became numbers
Words such as attention, wellbeing, misinformation and engagement are not measurements until operationalized. Was the outcome a validated instrument, device reading, diagnosis, self-report, administrative code or researcher judgment? Was the assessor blinded to group? Did the threshold exist before analysis?
Check whether the paper measured the outcome readers care about. Minutes in an app do not necessarily equal learning. Clicks do not equal persuasion. A proxy can be useful, but the inference from proxy to real-world claim must remain visible.
5. Find the estimate before the verdict
“Statistically significant” is not an effect size. Find the point estimate, confidence interval and units. For risks, translate relative effects into absolute terms using a relevant baseline. A 50 percent relative reduction means something different when risk falls from 20 in 100 to 10 in 100 than when it falls from 2 in 10,000 to 1 in 10,000.
Cochrane guidance emphasizes interpreting estimates with confidence intervals and avoiding a bright-line split between significant and non-significant. A small p-value does not tell you the effect is important. A p-value above a threshold does not prove no effect. Ask whether the interval includes benefits, harms or trivial changes large enough to alter the decision.
6. Separate imprecision from bias
A wide interval is visible uncertainty. Bias can move the entire estimate while leaving the interval narrow. Look for allocation problems, unblinded outcome assessment, selective reporting, deviations from intended treatment, uncontrolled confounding, multiple analyses and outcomes disclosed only after results were known.
Reporting guidelines help reveal what should be present. CONSORT provides a framework for randomized trials; STROBE covers major observational designs. Their purpose is better reporting, not a quality stamp. A paper can follow a checklist and still have a weak question or biased design. Use the checklist to locate information, then judge the information.
7. Compare the paper with the plan
When a protocol, trial registration or preregistration exists, compare primary outcomes, exclusions, sample size and analysis. Changes are not automatically misconduct; real research encounters surprises. Undisclosed flexibility is the problem because it lets results determine which question appears to have been asked all along.
Record whether the outcome was prespecified, changed with explanation or impossible to verify. Do not treat registration as proof. A registered bad design remains bad, and fields differ in how complete or common registration is.
8. Place one paper inside the evidence
A paper rarely begins a subject. Read what earlier evidence it cites, then search for systematic reviews, replications and serious contrary results. Check whether the new study resolves a known limitation or simply produces a more shareable estimate. One surprising result can be valuable without overturning a mature body of evidence.
Use the Reality Audit to trace the public claim back to the paper, and protect the original question with Search Before Wonder. The goal is not to outsource judgment to the first review you find. It is to learn whether the paper is typical, corrective or isolated.
9. Read funding, conflicts and corrections
Funding does not invalidate a result, and a conflict declaration does not repair a method. Treat both as information about where independent replication and scrutiny matter. Look for author affiliations, sponsor roles in design and analysis, data availability, corrections, expressions of concern and retractions.
Also inspect your own conflict: do you want the paper to win an argument? The feeling of understanding becomes especially persuasive when a result confirms identity or prior belief.
10. Write a three-layer note
End with three headings. Under reported, write what the study directly found. Under inferred, write the broader interpretation you think the evidence supports. Under open, write the uncertainties, competing explanations and information needed to update.
Then add one sentence: “This would change my decision if…” If you cannot connect the paper to any decision or model, you may be collecting conclusions rather than learning.
The twelve-question reading card
- What exact question was asked?
- Which design created the comparison?
- Who entered, who was excluded and who left?
- How were exposure and outcome measured?
- Was the outcome important or merely convenient?
- What is the effect in absolute as well as relative terms?
- What range does the interval leave open?
- Which biases could move the estimate?
- Were outcomes and analyses specified in advance?
- Does this population resemble the one I care about?
- How does the result fit the wider evidence?
- What observation would change my interpretation?
Primary and authoritative reading tools
- Cochrane, Handbook Chapter 15: Interpreting results and drawing conclusions.
- Cochrane, Handbook Chapter 7: Considering bias and conflicts of interest.
- EQUATOR Network, CONSORT reporting guideline and STROBE reporting guideline.
- ClinicalTrials.gov, registration and reporting requirements.
END OF FIELD GUIDE 032
Keep the question. Test the model.
Choose the narrowest claim the evidence can carry, then leave room for revision.