Statistics connects a question, a target population, measurements, a study design, an analysis, and uncertainty. A calculation can be arithmetically correct while the conclusion is unsupported.
Population, sample, parameter, statistic
The target population is the set of units and period addressed by the question. A sample is the observed subset selected by a documented mechanism. A parameter is an unknown population quantity; a statistic is calculated from sample data. A sample supports generalization only to the extent that coverage, selection, measurement, response, weighting, and analysis support it.
Describe location and spread
- Mean: the sum divided by the number of observations; useful for additive quantities but sensitive to extremes.
- Median: the middle ordered value; resistant to extreme magnitudes.
- Mode: the most frequent value or category; there may be none or several.
- Variance: the average squared deviation under the stated population or sample convention. Population variance uses
N; the usual sample variance usesn-1. - Standard deviation: the square root of variance, expressed in the variable's units.
No summary is universally best. Report the distribution, missingness, denominator, and units; plots and robust spread measures can expose structure hidden by one number.
Probability models are conditional
A normal distribution is a model, not a default law for data. The familiar 68–95–99.7 percentages apply to an ideal normal distribution. Real measurements can be skewed, discrete, bounded, multimodal, dependent, or heavy-tailed. Check whether an approximation is suitable for the task.
Estimation and intervals
A point estimate should be accompanied by an uncertainty interval when the design permits it. A 95% confidence interval is produced by a procedure that would cover the fixed parameter in 95% of repeated samples under its assumptions; it is not automatically a 95% probability that this realized interval contains the parameter.
Hypothesis tests without myths
A p-value is the probability, under the statistical model and null hypothesis, of a result at least as incompatible with that model as the observed result. It is not the probability that the null is true, the probability results occurred by chance, or a measure of practical importance.
Specify the estimand, test, assumptions, alpha, direction, multiplicity plan, exclusions, and stopping rule before inspecting outcomes when possible. Report effect estimates and compatible intervals. A threshold such as 0.05 is a convention, not a universal decision rule; failing to reject does not prove equivalence or no effect.
Correlation and regression
Correlation summarizes association under a chosen measure; it does not establish causation. Linear regression estimates a conditional relationship under assumptions about functional form and errors. Inspect residuals, influential observations, dependence, leakage, temporal stability, and out-of-sample performance. Estimating the causal effect of an intervention requires a defensible design or causal assumptions.
A practical analysis sequence
- Define the decision, population, outcome, exposure, and time window.
- Document sampling or assignment and measurement.
- Inspect missingness and distributions without hiding exclusions.
- Choose an estimand and method whose assumptions fit the design.
- Report effect size, uncertainty, sensitivity checks, and limits.
Continue with examples of bad data visualization, feature selection techniques, and the bias–variance tradeoff.

Historical comments from Datanizant
No public comments on this article
No approved public comments were included in the WordPress export for this article.