Does the study prove the headline?
A result can be real and still be too small, too indirect or too specific to support the claim placed above it.
“Associated with” and “caused by” are different sentences.
The study design, comparison, population and outcome decide which sentence is defensible. A prestigious journal cannot repair the wrong design for the question.
Two things vary together.
Active people may have lower disease rates. Other differences between them may explain part of the pattern.
Changing A changes B.
This needs a design that rules out credible alternatives. Randomisation helps, but execution, dropout and measurement still matter.
The result fits this person and choice.
Evidence can be internally valid and still apply poorly to another age, condition, setting, dose or outcome.
Build the study behind the headline.
Change the design and outcome. The same topic can support very different claims.
The headline outruns the design.
A snapshot can show that exposure and outcome differ together at one time.
It cannot establish which came first or rule out confounding.
In this sample, activity level was associated with the measured marker.
Ask whether the outcome is the thing people actually care about.
A transient rise in BDNF is not the same as better examination performance. A one-point pain change is not automatically meaningful function. A stronger grip is not proof of added years of life.
Each design trades control for realism, scale or time.
Select a design. The hierarchy changes with the question: a trial is strong for an intervention, while a large cohort may be necessary to study rare outcomes over decades.
Mechanistic work asks: could this pathway operate?
Cells, animals, tissue, imaging or tightly controlled human experiments can reveal timing and biological pathways. Control may be high, while translation to daily life is uncertain.
- Strong for
- Mechanism, feasibility and precise measurement.
- Weak for
- Long-term patient outcomes and population-wide effect size.
- Check
- Was the model human, animal or cellular? Was the measured pathway necessary, sufficient or merely correlated?
Groups rarely differ in one thing only.
Observational research may compare people who chose different lives long before the study began.
Two groups, one visible exposure.
- More exercise
- Often better baseline health
- May sleep differently
- May have higher income
- Less exercise
- May include early disease
- Different work demands
- Different care access
If Group A later has less dementia, exercise may contribute. Baseline health, education, income, sleep or early symptoms may also contribute.
Statistical adjustment helps. It does not randomise history.
0 / 5 safeguards
The crude association is vulnerable to obvious alternative explanations.
Early disease can reduce activity before diagnosis.
If lower activity appears before dementia, inactivity might contribute to risk; subtle prodromal disease might also have reduced activity. Excluding diagnoses soon after baseline and repeating activity measurement can reduce this problem, not erase it.
Read the method before falling in love with the result.
Tap each pass. You can use the sequence on a paper, podcast claim or newspaper headline.
What exact question was the study designed to answer?
Translate the aim into population, exposure or intervention, comparison and outcome. If the headline asks about dementia but the study measured reaction time after one workout, name the gap.
Ask: Is this a mechanism, prediction, treatment or prevention question?
Slow the jump from result to recommendation.
It helps identify design, population, comparison, outcome, uncertainty and the language the data can support.
Replace expert appraisal or patient-specific judgement.
Bias tools, statistics, clinical context and the full body of evidence still matter. A checklist cannot turn a poor paper into a useful one.
Built from evidence-appraisal principles.
- Cochrane. Cochrane Handbook for Systematic Reviews of Interventions.
- Sterne JAC et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. Cochrane, 2019.
- Sterne JA et al. ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions. BMJ. 2016;355:i4919.
- Guyatt GH et al. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ. 2008;336:924–926.
- Hernán MA, Robins JM. Causal Inference: What If. Chapman & Hall/CRC, 2020.