Fact-check note: Reviewed September 4, 2026. This applied guide complements the canonical concepts article and removes universal threshold claims.
A p-value is the probability, under a specified statistical model and null hypothesis, of obtaining a result at least as incompatible with that hypothesis as the observed result. It is not the probability that the null is true, that results arose “by chance,” or that an effect is real.
Read the estimate first
Report the effect estimate, units, target population, comparison, study design, confidence interval, p-value where useful, and assumptions. Statistical significance does not establish importance, causality, absence of bias, or replicability. A large, precise trivial effect and a practically important but imprecise effect require different decisions.
Interpret confidence intervals carefully
A 95% frequentist confidence procedure produces intervals that contain the fixed target in 95% of repeated samples under the model and procedure. For one computed interval, do not say there is a 95% probability the fixed parameter lies inside. Values inside are more compatible with the data and assumptions than values outside, but the interval is not a list of equally plausible values.
Connect intervals and tests
For matching two-sided procedures, a 95% interval excluding the null corresponds to a p-value below 0.05. Mismatches arise with one-sided tests, different models, corrections, or rounding. Predefine the estimand, clinically or operationally meaningful effect, alpha, sample-size method, and analysis rather than selecting thresholds after seeing results.
Account for multiplicity
Testing many hypotheses increases false-positive opportunities. With 20 independent tests, all nulls true, and alpha 0.05, the chance of at least one rejection is 1 − 0.9520, about 64.2%. Dependence changes this value. Choose family-wise-error, false-discovery-rate, hierarchical, or selective-inference methods based on the question; report the family and all analyses.
Use a reporting checklist
- State design, population, exclusions, outcomes, and analysis plan.
- Report estimates and uncertainty, not “significant/non-significant” alone.
- Describe model checks, missing data, multiplicity, and sensitivity analyses.
- Distinguish association from causal effects and preregister confirmatory work where feasible.
- Discuss practical thresholds, harms, costs, and external validity.
Review foundations in confidence and significance levels, avoid modeling errors with spurious correlation examples, and plan features using feature selection techniques.

Historical comments from Datanizant
No public comments on this article
No approved public comments were included in the WordPress export for this article.