Educational hub
How Science Works
How measurement, controls, causal inference, replication, uncertainty and evidence synthesis constrain what conclusions the evidence can support.
Science is not a list of approved answers. It is a collection of methods for making error easier to detect.
That distinction matters for Claim by Claim. We do not ask whether a statement agrees with an institution. We ask what observations would be expected if the claim were true, how those observations were measured, what alternative explanations remain, and how strong the total evidence is.
A claim has to become testable
“Energy is blocked” is not yet a scientific claim.
A useful empirical claim identifies something measurable:
- what changes;
- compared with what;
- under which conditions;
- by how much;
- over what time period;
- and what result would count against the claim.
This is why operational definitions matter. A vague proposition can survive almost any result. A precise proposition risks being wrong—and that is a feature, not a flaw.
Mechanism and outcome are different questions
A proposed mechanism can be plausible while a claimed outcome is unsupported. An outcome can also appear in data before the mechanism is understood.
We therefore separate questions such as:
Can the proposed physical or biological process occur?
from:
Does it occur at the relevant dose or exposure?
and from:
Does it produce the claimed real-world benefit or harm?
Collapsing those steps is a common source of overclaiming. Detecting an effect in cells, for example, does not automatically establish a meaningful effect in humans at ordinary exposure levels.
Controls help distinguish the claim from alternatives
An observation by itself rarely identifies its cause.
If symptoms improve after a treatment, possible explanations include the treatment, natural recovery, regression toward the mean, concurrent treatment, expectation, measurement variability or selective reporting.
Controls and comparison groups are tools for separating those possibilities.
The right control depends on the question. A placebo control can address expectation effects. A sham procedure can help with device or procedural interventions. A matched comparison group may be necessary when randomization is impossible. Historical controls are weaker for many causal questions because conditions can change over time.
Correlation does not automatically establish causation
When two variables move together, several explanations remain possible:
A causes B.
B causes A.
A third variable affects both.
The association is partly or entirely due to selection, measurement or chance.
Causal inference is therefore a design problem as much as a statistical problem. Randomization can help balance known and unknown confounders. Natural experiments, longitudinal designs and careful adjustment can strengthen observational inference, but they do not all provide the same level of causal identification.
Statistical significance is not effect size
A small p-value does not tell you whether an effect is large, useful or important.
With enough data, a tiny effect can be statistically detectable. With too little data, an important effect can remain statistically uncertain.
That is why we care about effect size, confidence intervals, absolute risk, baseline risk and measurement quality, not just whether a result crossed a conventional significance threshold.
For a reader, the useful question is often not “Was p < 0.05?” but “How large is the estimated effect, how uncertain is that estimate, and would a difference of that size matter?”
One study is evidence, not the end of the question
Individual studies are vulnerable to chance, design limitations, analytic flexibility, measurement error and context-specific effects.
Replication asks whether a finding appears again under sufficiently similar conditions. Reproducibility asks, in one common usage, whether the same results can be obtained from the same data and computational procedures.
Failures to replicate do not all mean fraud or useless science. They can expose hidden moderators, weak measurements, low power, selective publication or genuine context dependence.
The scientific advantage is that the system has a mechanism for discovering those problems.
Evidence synthesis is stronger than citation counting
Five papers do not automatically outweigh two papers.
The relevant questions include:
- how large were the studies;
- what designs did they use;
- how directly did they test the claim;
- what was the risk of bias;
- were results consistent;
- are the studies independent;
- is there evidence of missing negative results;
- and how applicable are the populations, doses and outcomes to the exact claim?
Systematic reviews and meta-analyses are designed to synthesize evidence more systematically, but they inherit weaknesses from the studies they include. A meta-analysis of biased or highly heterogeneous studies does not magically become high-quality evidence.
Evidence quality and evidence status are not the same thing
A conclusion can be unsupported because the available evidence is weak or sparse.
A conclusion can be contradicted because strong evidence points the other way.
Those are not interchangeable.
Likewise, a high-quality body of evidence can support a small effect, while a low-quality body of evidence can report a dramatic effect with substantial uncertainty.
Frameworks such as GRADE formalize this separation by asking how certain we are in an effect estimate before moving to recommendations or decisions.
Consensus is evidence about the evidence, not an oracle
Scientific consensus is useful because it can summarize a large distributed process: many studies, methods, failed attempts, replications, technical disagreements and accumulated domain knowledge.
But consensus is not infallible and it is not a substitute for showing the evidence trail.
For Claim by Claim, an institutional statement can be useful context, especially when it synthesizes a mature field. It does not receive automatic priority over stronger contradictory primary evidence simply because of the institution’s name.
The same standard runs in both directions: institutional distrust is not evidence of error, and institutional authority is not evidence of truth by itself.
Uncertainty is information
A scientifically honest answer is often a range rather than a binary verdict.
Uncertainty can come from sampling error, measurement error, model assumptions, indirectness, heterogeneity, missing data, confounding or simply an immature evidence base.
Good evidence communication makes those limitations visible instead of hiding them behind certainty language.
That is why our evidence statuses distinguish supported, mostly supported, misleading, unsupported, contradicted, inconclusive and unfalsifiable rather than forcing every investigation into true/false.
The central question: what would change the conclusion?
Science works because claims can lose.
A useful empirical proposition should identify observations that would reduce confidence in it. If every possible outcome can be reinterpreted as confirmation, the claim has moved outside normal empirical testing.
This is the logic behind the recurring section What Would Change Our Conclusion? in every Claim by Claim investigation.
It is not a rhetorical flourish. It is the mechanism that keeps the conclusion conditional on evidence rather than identity, authority or prior commitment.
Research trail
Sources & further reading
These sources support the mechanisms and methodological points discussed in this guide. They are not a claim that every finding applies identically to every person or context.
- Reproducibility and Replicability in Science National Academies of Sciences, Engineering, and Medicine · 2019 · consensus statement · DOI 10.17226/25303
Consensus report defining reproducibility and replicability and reviewing why failures to reproduce or replicate can occur.
- Cochrane Handbook for Systematic Reviews of Interventions Cochrane · 2024 · standard
Methodological handbook covering study selection, risk of bias, effect measures, meta-analysis, missing results, GRADE and interpretation of evidence.
- GRADE: an emerging consensus on rating quality of evidence and strength of recommendations BMJ · 2008 · academic resource · DOI 10.1136/bmj.39489.470347.AD
Introduces the GRADE framework for separating certainty of evidence from the strength of recommendations.
- Estimating the reproducibility of psychological science Science · 2015 · peer reviewed study · DOI 10.1126/science.aac4716
Large collaborative replication project illustrating why individual published findings should not be treated as the final word and why replication is informative.
- The ASA's Statement on p-Values: Context, Process, and Purpose The American Statistician · 2016 · consensus statement · DOI 10.1080/00031305.2016.1154108
American Statistical Association statement explaining why p-values do not measure effect size, importance or the probability that a hypothesis is true.
Connected reasoning patterns