A confidence level describes the long-run coverage of a specified interval procedure under its assumptions. A significance level, α, is a test procedure's pre-specified upper bound on the probability of rejecting the null when that null and the model assumptions hold. Neither is the probability that a particular conclusion is correct.

What confidence intervals tell you

A confidence interval combines an estimate with uncertainty calculated by a method appropriate to the design and estimand. “Estimate ± margin of error” is common but not universal; proportions, ratios, clusters, time series, complex surveys, and resampling can require different methods.

For a procedure with 95% confidence coverage, approximately 95% of intervals from repeated samples contain the fixed true parameter under the procedure's assumptions. After one frequentist interval is computed, 95% is not the posterior probability that the parameter lies inside that particular interval.

For the same data and method, higher confidence generally produces a wider interval. Width also depends on sample size, variability, design effects, and assumptions. A narrow interval can still mislead when sampling is biased or the estimand does not match the decision.

What significance levels and p-values tell you

Before confirmatory analysis, specify the null and alternative hypotheses, test statistic, assumptions, sidedness, and α. Under a simple null and calibrated procedure, α controls the long-run Type I error rate. It is not the probability that this rejection is a mistake or that the null is true.

A p-value is the probability, under the null model and test assumptions, of a test statistic at least as extreme as the observed one. A small p-value indicates incompatibility with that specified model; it does not say that chance caused the data, give a hypothesis probability, or measure effect size.

  • If p ≤ pre-specified α: reject the null under the rule; report the exact p-value, effect estimate, interval, assumptions, and analysis choices.
  • If p > α: fail to reject the null. This is not acceptance of the null or proof of no effect.

How matching tests and intervals connect

When a confidence interval is created by inverting a particular hypothesis test, a two-sided level-α test corresponds to a two-sided 100(1 − α)% interval under the same model, parameterization, standard error, tail convention, and assumptions. The null value is rejected exactly when it lies outside that matched interval.

This is not an identity between any reported interval and any test. One-sided, bootstrap, Bayesian, noninferiority, equivalence, sequential, adjusted, and multiple-comparison procedures have different relationships.

A matched example

Suppose a pre-specified two-sided level-0.05 test compares conversion-rate differences and its matching 95% interval is +0.5 to +4.5 percentage points. Excluding zero corresponds to rejection at 0.05 under the shared assumptions. The interval does not guarantee that a future rollout will achieve that lift.

Statistical significance is not a launch decision

A valid result with p = 0.02 and α = 0.05 crosses the specified threshold. Deployment still requires the estimated effect and interval, minimum practically important effect, randomization and exposure checks, costs, guardrails, segment behavior, missing outcomes, novelty effects, and multiplicity or repeated-peeking correction.

A 0.01% relative change differs from 0.01 percentage points. Whether either matters depends on the baseline rate, traffic, costs, benefits, duration, uncertainty, and risks. State the scale and estimate decision value. For broader decision context, see data-driven decision-making.

Multiplicity, power, and sample size

Trying many outcomes, segments, models, stopping times, or exclusions and reporting only successes inflates false-positive risk. Pre-specify confirmatory analyses where appropriate, distinguish exploration from confirmation, adjust the analysis family when required, and report material choices. Do not treat p = 0.049 and p = 0.051 as opposite truths.

Power depends on the alternative effect, variability, sample size, test, design, and α. Lowering α while holding other inputs fixed generally lowers power. Larger samples often reduce standard error and increase power for a specified nonzero effect, but do not repair bias, confounding, dependence, measurement error, informative missingness, or misspecification.

One-sided tests

A one-sided test can be appropriate when its direction, null, decision consequences, and plan are justified before seeing data. Do not choose it after viewing the estimate. Use a matching one-sided bound and explain the decision rule; an adverse result in the opposite direction remains important.

Reporting checklist

  • Name the estimand, population, sampling or randomization design, and sample sizes.
  • Report the effect scale, exact p-value, interval method and level, sidedness, and assumptions.
  • Disclose exclusions, missing-data handling, sequential looks, and multiplicity.
  • Separate statistical, practical, and causal conclusions.
  • Preserve an untouched test set in predictive modeling; see leakage-safe model selection.

Uncertainty also affects model complexity decisions; the bias–variance trade-off provides complementary context, while spurious-correlation examples illustrate why significance alone is insufficient.

Frequently asked questions

Is 95% confidence always the right choice?

No. Choose the procedure and level for the domain, decision costs, design, and reporting standard before inspecting results.

Does p > 0.05 prove no effect?

No. It means the test did not reject at that threshold. Inspect the estimate, interval, design, power, and compatibility with effects that matter.

Does p < 0.05 prove causation?

No. Causal interpretation depends on design, identification assumptions, execution, and bias—not the threshold alone.