Free tool 04 · Evidence literacy

Does the study prove the headline?

A result can be real and still be too small, too indirect or too specific to support the claim placed above it.

Built by ZACH · Reviewed 30 September 2026 · Educational, not a medical evidence-grading service

Start with language

“Associated with” and “caused by” are different sentences.

The study design, comparison, population and outcome decide which sentence is defensible. A prestigious journal cannot repair the wrong design for the question.

Association

Two things vary together.

Active people may have lower disease rates. Other differences between them may explain part of the pattern.

Causation

Changing A changes B.

This needs a design that rules out credible alternatives. Randomisation helps, but execution, dropout and measurement still matter.

Application

The result fits this person and choice.

Evidence can be internally valid and still apply poorly to another age, condition, setting, dose or outcome.

01 / Interactive claim lab

Build the study behind the headline.

Change the design and outcome. The same topic can support very different claims.

Current verdict
D

The headline outruns the design.

What the study may support

A snapshot can show that exposure and outcome differ together at one time.

What it cannot support

It cannot establish which came first or rule out confounding.

Rewrite the headline

In this sample, activity level was associated with the measured marker.

The result is one link

Ask whether the outcome is the thing people actually care about.

A transient rise in BDNF is not the same as better examination performance. A one-point pain change is not automatically meaningful function. A stronger grip is not proof of added years of life.

02 / Study designs

Each design trades control for realism, scale or time.

Select a design. The hierarchy changes with the question: a trial is strong for an intervention, while a large cohort may be necessary to study rare outcomes over decades.

Mechanistic work asks: could this pathway operate?

Cells, animals, tissue, imaging or tightly controlled human experiments can reveal timing and biological pathways. Control may be high, while translation to daily life is uncertain.

Strong for
Mechanism, feasibility and precise measurement.
Weak for
Long-term patient outcomes and population-wide effect size.
Check
Was the model human, animal or cellular? Was the measured pathway necessary, sufficient or merely correlated?
Typical causal confidence for an intervention question
03 / Confounding

Groups rarely differ in one thing only.

Observational research may compare people who chose different lives long before the study began.

Two groups, one visible exposure.

More active Group A
  • More exercise
  • Often better baseline health
  • May sleep differently
  • May have higher income
Less active Group B
  • Less exercise
  • May include early disease
  • Different work demands
  • Different care access

If Group A later has less dementia, exercise may contribute. Baseline health, education, income, sleep or early symptoms may also contribute.

Statistical adjustment helps. It does not randomise history.

0 / 5 safeguards

The crude association is vulnerable to obvious alternative explanations.

Reverse causation

Early disease can reduce activity before diagnosis.

If lower activity appears before dementia, inactivity might contribute to risk; subtle prodromal disease might also have reduced activity. Excluding diagnoses soon after baseline and repeating activity measurement can reduce this problem, not erase it.

04 / Six-pass paper reader

Read the method before falling in love with the result.

Tap each pass. You can use the sequence on a paper, podcast claim or newspaper headline.

What exact question was the study designed to answer?

Translate the aim into population, exposure or intervention, comparison and outcome. If the headline asks about dementia but the study measured reaction time after one workout, name the gap.

Ask: Is this a mechanism, prediction, treatment or prevention question?

What this tool can do

Slow the jump from result to recommendation.

It helps identify design, population, comparison, outcome, uncertainty and the language the data can support.

What it cannot do

Replace expert appraisal or patient-specific judgement.

Bias tools, statistics, clinical context and the full body of evidence still matter. A checklist cannot turn a poor paper into a useful one.

Sources & methods

Built from evidence-appraisal principles.

  1. Cochrane. Cochrane Handbook for Systematic Reviews of Interventions.
  2. Sterne JAC et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. Cochrane, 2019.
  3. Sterne JA et al. ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions. BMJ. 2016;355:i4919.
  4. Guyatt GH et al. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ. 2008;336:924–926.
  5. Hernán MA, Robins JM. Causal Inference: What If. Chapman & Hall/CRC, 2020.

The letter grade in the claim lab is a teaching prompt created for this page; it is not the GRADE framework and must not be reported as a formal risk-of-bias judgement.