Why a Correlation Can Mislead
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This example demonstrates multiple-comparison and selection effects: searching many series makes extreme sample correlations more likely even when variables are unrelated. “P-hacking” is broader and should not be used as a synonym for every cherry-picked correlation. Report the search process, number of comparisons, uncertainty, and out-of-sample replication.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
The seven examples below are not all empirical studies. Some are sourced demonstrations, some are satire built from public series, and some are teaching hypotheticals. For each, ask who produced the data, how variables were defined, whether the window was selected in advance, how many comparisons were searched, what causal graph is plausible, and whether the relationship persists in new data.
1. Ice Cream Sales and Drowning Deaths

- Always consider seasonal factors when analyzing time-series data. Seasonal variations can create misleading correlations between variables.
- Look for third variables that might explain the observed correlation. Conduct thorough exploratory data analysis to identify potential confounding variables.
- Consider time-based confounding variables. Like seasonality, other time-related factors can create spurious relationships.
- Use this classic example to teach and reinforce the difference between correlation and causation. It's a simple yet powerful way to communicate this crucial statistical concept.
2. Divorce Rate in Maine vs Margarine Consumption

3. Number of Pirates vs Global Temperature
The pirates-versus-temperature graphic originated as satire associated with Bobby Henderson’s Flying Spaghetti Monster argument. It juxtaposes rough pirate-count estimates with temperature values to mock the leap from temporal association to causation. It is not a validated historical dataset or climate-attribution analysis.
Tyler Vigen published a deliberately selected annual series pairing Nicolas Cage film appearances with swimming-pool drownings. Treat the displayed coefficient and date range as properties of that selected sample. Extending the window, changing definitions, or testing new data can change the relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This example demonstrates multiple-comparison and selection effects: searching many series makes extreme sample correlations more likely even when variables are unrelated. “P-hacking” is broader and should not be used as a synonym for every cherry-picked correlation. Report the search process, number of comparisons, uncertainty, and out-of-sample replication.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
The seven examples below are not all empirical studies. Some are sourced demonstrations, some are satire built from public series, and some are teaching hypotheticals. For each, ask who produced the data, how variables were defined, whether the window was selected in advance, how many comparisons were searched, what causal graph is plausible, and whether the relationship persists in new data.
1. Ice Cream Sales and Drowning Deaths

- Always consider seasonal factors when analyzing time-series data. Seasonal variations can create misleading correlations between variables.
- Look for third variables that might explain the observed correlation. Conduct thorough exploratory data analysis to identify potential confounding variables.
- Consider time-based confounding variables. Like seasonality, other time-related factors can create spurious relationships.
- Use this classic example to teach and reinforce the difference between correlation and causation. It's a simple yet powerful way to communicate this crucial statistical concept.
2. Divorce Rate in Maine vs Margarine Consumption

3. Number of Pirates vs Global Temperature
The pirates-versus-temperature graphic originated as satire associated with Bobby Henderson’s Flying Spaghetti Monster argument. It juxtaposes rough pirate-count estimates with temperature values to mock the leap from temporal association to causation. It is not a validated historical dataset or climate-attribution analysis.
Use it only as a rhetorical illustration. Evidence about climate causes comes from physical measurements, attribution methods, and converging climate science—not from accepting or rejecting this cartoon correlation.
4. Nicolas Cage Movies and Swimming Pool Drownings

- Selected annual-series correlation: The close tracking of the two variables over a significant period provides a striking visual representation of a spurious correlation.
- Involves Entertainment Industry and Safety Statistics: The juxtaposition of these seemingly unrelated fields further highlights the randomness of the correlation.
- Demonstrates Coincidental Trending Patterns: This example perfectly illustrates how unrelated trends can align by chance.
- Easy to Visualize and Understand: The simplicity of the comparison makes it readily accessible to a broad audience, including those without a strong statistical background.
- Entertaining Way to Teach Statistical Concepts: The inherent humor of the example makes it an engaging way to introduce the concept of spurious correlations.
- Shows Absurdity of Assuming Causation from Correlation: This example effectively dismantles the common misconception that correlation implies causation.
- Memorable Celebrity Connection Aids Retention: The inclusion of Nicolas Cage adds a memorable element that strengthens the learning experience.
- May Trivialize Actual Drowning Statistics: Using a serious topic like accidental drownings in a humorous context could be perceived as insensitive.
- Represents Selective Data Mining: It highlights the potential for misleading results when data is selectively chosen to support a particular narrative.
- Could be Misunderstood as Intentionally Meaningful: Some individuals might misinterpret the correlation as having some underlying, albeit bizarre, meaning.
- Question correlations involving celebrities or pop culture: These are often ripe for spurious relationships due to the high visibility and frequent media coverage of these figures.
- Consider sample sizes and time periods in correlation analysis: Limited data sets and specific timeframes can create misleading correlations.
- Look for logical explanatory mechanisms: If a correlation lacks a plausible explanation, it should be treated with skepticism.
- Use entertaining examples like this to engage students while teaching serious concepts: This example provides a valuable lesson in critical thinking and data analysis.
5. Chocolate Consumption and Nobel Prize Winners
In a 2012 New England Journal of Medicine correspondence, Franz Messerli compared national chocolate consumption with Nobel laureates per capita and reported a country-level correlation. It did not establish that eating chocolate causes Nobel-level achievement. Country-level relationships cannot be transferred to individuals, and observational cross-country comparisons are vulnerable to confounding, measurement differences, and influential observations.
National income, education and research investment, population history, data quality, and other country-level factors are plausible alternative explanations, but this small ecological comparison does not identify which factor explains the association. Publication venue does not turn an ecological correlation into a causal estimate.
6. Autism Identification and Organic-Food Sales
A widely circulated teaching graph places two U.S. time series on the same chart: reported autism identification and organic-food sales. Both rise over part of the selected period. That visual alignment does not demonstrate a causal relationship, and this article found no credible evidence that organic-food consumption causes autism.
This is primarily a shared-trend problem. A responsible analysis would verify both source series, align definitions and units, disclose selected years, examine changes rather than levels, account for autocorrelation, and test a causal model before making any health claim.
CDC cautions that changes in measured autism prevalence reflect multiple factors, including diagnostic practices, screening, awareness, access to services, population coverage, and potentially other influences. Use precise language such as “identified prevalence” or “reported prevalence.”
Shoe size and reading ability is a standard classroom example, often demonstrated with toy or classroom data. Across a mixed-age sample of children, both measures can increase with age and schooling, creating an association that should not be interpreted as an effect of shoe size on reading.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This example demonstrates multiple-comparison and selection effects: searching many series makes extreme sample correlations more likely even when variables are unrelated. “P-hacking” is broader and should not be used as a synonym for every cherry-picked correlation. Report the search process, number of comparisons, uncertainty, and out-of-sample replication.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
The seven examples below are not all empirical studies. Some are sourced demonstrations, some are satire built from public series, and some are teaching hypotheticals. For each, ask who produced the data, how variables were defined, whether the window was selected in advance, how many comparisons were searched, what causal graph is plausible, and whether the relationship persists in new data.
1. Ice Cream Sales and Drowning Deaths

- Always consider seasonal factors when analyzing time-series data. Seasonal variations can create misleading correlations between variables.
- Look for third variables that might explain the observed correlation. Conduct thorough exploratory data analysis to identify potential confounding variables.
- Consider time-based confounding variables. Like seasonality, other time-related factors can create spurious relationships.
- Use this classic example to teach and reinforce the difference between correlation and causation. It's a simple yet powerful way to communicate this crucial statistical concept.
2. Divorce Rate in Maine vs Margarine Consumption

3. Number of Pirates vs Global Temperature
The pirates-versus-temperature graphic originated as satire associated with Bobby Henderson’s Flying Spaghetti Monster argument. It juxtaposes rough pirate-count estimates with temperature values to mock the leap from temporal association to causation. It is not a validated historical dataset or climate-attribution analysis.
Tyler Vigen published a deliberately selected annual series pairing Nicolas Cage film appearances with swimming-pool drownings. Treat the displayed coefficient and date range as properties of that selected sample. Extending the window, changing definitions, or testing new data can change the relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This example demonstrates multiple-comparison and selection effects: searching many series makes extreme sample correlations more likely even when variables are unrelated. “P-hacking” is broader and should not be used as a synonym for every cherry-picked correlation. Report the search process, number of comparisons, uncertainty, and out-of-sample replication.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
The seven examples below are not all empirical studies. Some are sourced demonstrations, some are satire built from public series, and some are teaching hypotheticals. For each, ask who produced the data, how variables were defined, whether the window was selected in advance, how many comparisons were searched, what causal graph is plausible, and whether the relationship persists in new data.
1. Ice Cream Sales and Drowning Deaths

- Always consider seasonal factors when analyzing time-series data. Seasonal variations can create misleading correlations between variables.
- Look for third variables that might explain the observed correlation. Conduct thorough exploratory data analysis to identify potential confounding variables.
- Consider time-based confounding variables. Like seasonality, other time-related factors can create spurious relationships.
- Use this classic example to teach and reinforce the difference between correlation and causation. It's a simple yet powerful way to communicate this crucial statistical concept.
2. Divorce Rate in Maine vs Margarine Consumption

3. Number of Pirates vs Global Temperature
The pirates-versus-temperature graphic originated as satire associated with Bobby Henderson’s Flying Spaghetti Monster argument. It juxtaposes rough pirate-count estimates with temperature values to mock the leap from temporal association to causation. It is not a validated historical dataset or climate-attribution analysis.
Use it only as a rhetorical illustration. Evidence about climate causes comes from physical measurements, attribution methods, and converging climate science—not from accepting or rejecting this cartoon correlation.
4. Nicolas Cage Movies and Swimming Pool Drownings

- Selected annual-series correlation: The close tracking of the two variables over a significant period provides a striking visual representation of a spurious correlation.
- Involves Entertainment Industry and Safety Statistics: The juxtaposition of these seemingly unrelated fields further highlights the randomness of the correlation.
- Demonstrates Coincidental Trending Patterns: This example perfectly illustrates how unrelated trends can align by chance.
- Easy to Visualize and Understand: The simplicity of the comparison makes it readily accessible to a broad audience, including those without a strong statistical background.
- Entertaining Way to Teach Statistical Concepts: The inherent humor of the example makes it an engaging way to introduce the concept of spurious correlations.
- Shows Absurdity of Assuming Causation from Correlation: This example effectively dismantles the common misconception that correlation implies causation.
- Memorable Celebrity Connection Aids Retention: The inclusion of Nicolas Cage adds a memorable element that strengthens the learning experience.
- May Trivialize Actual Drowning Statistics: Using a serious topic like accidental drownings in a humorous context could be perceived as insensitive.
- Represents Selective Data Mining: It highlights the potential for misleading results when data is selectively chosen to support a particular narrative.
- Could be Misunderstood as Intentionally Meaningful: Some individuals might misinterpret the correlation as having some underlying, albeit bizarre, meaning.
- Question correlations involving celebrities or pop culture: These are often ripe for spurious relationships due to the high visibility and frequent media coverage of these figures.
- Consider sample sizes and time periods in correlation analysis: Limited data sets and specific timeframes can create misleading correlations.
- Look for logical explanatory mechanisms: If a correlation lacks a plausible explanation, it should be treated with skepticism.
- Use entertaining examples like this to engage students while teaching serious concepts: This example provides a valuable lesson in critical thinking and data analysis.
5. Chocolate Consumption and Nobel Prize Winners
In a 2012 New England Journal of Medicine correspondence, Franz Messerli compared national chocolate consumption with Nobel laureates per capita and reported a country-level correlation. It did not establish that eating chocolate causes Nobel-level achievement. Country-level relationships cannot be transferred to individuals, and observational cross-country comparisons are vulnerable to confounding, measurement differences, and influential observations.
National income, education and research investment, population history, data quality, and other country-level factors are plausible alternative explanations, but this small ecological comparison does not identify which factor explains the association. Publication venue does not turn an ecological correlation into a causal estimate.
6. Autism Identification and Organic-Food Sales
A widely circulated teaching graph places two U.S. time series on the same chart: reported autism identification and organic-food sales. Both rise over part of the selected period. That visual alignment does not demonstrate a causal relationship, and this article found no credible evidence that organic-food consumption causes autism.
This is primarily a shared-trend problem. A responsible analysis would verify both source series, align definitions and units, disclose selected years, examine changes rather than levels, account for autocorrelation, and test a causal model before making any health claim.
CDC cautions that changes in measured autism prevalence reflect multiple factors, including diagnostic practices, screening, awareness, access to services, population coverage, and potentially other influences. Use precise language such as “identified prevalence” or “reported prevalence.”
Adjusting or stratifying by age and educational stage can reduce or remove the association in the teaching example, but that result must be shown in the actual dataset rather than asserted universally. Conditioning should be justified by a causal model; mechanically controlling for everything can introduce bias.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This example demonstrates multiple-comparison and selection effects: searching many series makes extreme sample correlations more likely even when variables are unrelated. “P-hacking” is broader and should not be used as a synonym for every cherry-picked correlation. Report the search process, number of comparisons, uncertainty, and out-of-sample replication.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
The seven examples below are not all empirical studies. Some are sourced demonstrations, some are satire built from public series, and some are teaching hypotheticals. For each, ask who produced the data, how variables were defined, whether the window was selected in advance, how many comparisons were searched, what causal graph is plausible, and whether the relationship persists in new data.
1. Ice Cream Sales and Drowning Deaths

- Always consider seasonal factors when analyzing time-series data. Seasonal variations can create misleading correlations between variables.
- Look for third variables that might explain the observed correlation. Conduct thorough exploratory data analysis to identify potential confounding variables.
- Consider time-based confounding variables. Like seasonality, other time-related factors can create spurious relationships.
- Use this classic example to teach and reinforce the difference between correlation and causation. It's a simple yet powerful way to communicate this crucial statistical concept.
2. Divorce Rate in Maine vs Margarine Consumption

3. Number of Pirates vs Global Temperature
The pirates-versus-temperature graphic originated as satire associated with Bobby Henderson’s Flying Spaghetti Monster argument. It juxtaposes rough pirate-count estimates with temperature values to mock the leap from temporal association to causation. It is not a validated historical dataset or climate-attribution analysis.
Tyler Vigen published a deliberately selected annual series pairing Nicolas Cage film appearances with swimming-pool drownings. Treat the displayed coefficient and date range as properties of that selected sample. Extending the window, changing definitions, or testing new data can change the relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This example demonstrates multiple-comparison and selection effects: searching many series makes extreme sample correlations more likely even when variables are unrelated. “P-hacking” is broader and should not be used as a synonym for every cherry-picked correlation. Report the search process, number of comparisons, uncertainty, and out-of-sample replication.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
The seven examples below are not all empirical studies. Some are sourced demonstrations, some are satire built from public series, and some are teaching hypotheticals. For each, ask who produced the data, how variables were defined, whether the window was selected in advance, how many comparisons were searched, what causal graph is plausible, and whether the relationship persists in new data.
1. Ice Cream Sales and Drowning Deaths

- Always consider seasonal factors when analyzing time-series data. Seasonal variations can create misleading correlations between variables.
- Look for third variables that might explain the observed correlation. Conduct thorough exploratory data analysis to identify potential confounding variables.
- Consider time-based confounding variables. Like seasonality, other time-related factors can create spurious relationships.
- Use this classic example to teach and reinforce the difference between correlation and causation. It's a simple yet powerful way to communicate this crucial statistical concept.
2. Divorce Rate in Maine vs Margarine Consumption

3. Number of Pirates vs Global Temperature
The pirates-versus-temperature graphic originated as satire associated with Bobby Henderson’s Flying Spaghetti Monster argument. It juxtaposes rough pirate-count estimates with temperature values to mock the leap from temporal association to causation. It is not a validated historical dataset or climate-attribution analysis.
Use it only as a rhetorical illustration. Evidence about climate causes comes from physical measurements, attribution methods, and converging climate science—not from accepting or rejecting this cartoon correlation.
4. Nicolas Cage Movies and Swimming Pool Drownings

- Selected annual-series correlation: The close tracking of the two variables over a significant period provides a striking visual representation of a spurious correlation.
- Involves Entertainment Industry and Safety Statistics: The juxtaposition of these seemingly unrelated fields further highlights the randomness of the correlation.
- Demonstrates Coincidental Trending Patterns: This example perfectly illustrates how unrelated trends can align by chance.
- Easy to Visualize and Understand: The simplicity of the comparison makes it readily accessible to a broad audience, including those without a strong statistical background.
- Entertaining Way to Teach Statistical Concepts: The inherent humor of the example makes it an engaging way to introduce the concept of spurious correlations.
- Shows Absurdity of Assuming Causation from Correlation: This example effectively dismantles the common misconception that correlation implies causation.
- Memorable Celebrity Connection Aids Retention: The inclusion of Nicolas Cage adds a memorable element that strengthens the learning experience.
- May Trivialize Actual Drowning Statistics: Using a serious topic like accidental drownings in a humorous context could be perceived as insensitive.
- Represents Selective Data Mining: It highlights the potential for misleading results when data is selectively chosen to support a particular narrative.
- Could be Misunderstood as Intentionally Meaningful: Some individuals might misinterpret the correlation as having some underlying, albeit bizarre, meaning.
- Question correlations involving celebrities or pop culture: These are often ripe for spurious relationships due to the high visibility and frequent media coverage of these figures.
- Consider sample sizes and time periods in correlation analysis: Limited data sets and specific timeframes can create misleading correlations.
- Look for logical explanatory mechanisms: If a correlation lacks a plausible explanation, it should be treated with skepticism.
- Use entertaining examples like this to engage students while teaching serious concepts: This example provides a valuable lesson in critical thinking and data analysis.
5. Chocolate Consumption and Nobel Prize Winners
In a 2012 New England Journal of Medicine correspondence, Franz Messerli compared national chocolate consumption with Nobel laureates per capita and reported a country-level correlation. It did not establish that eating chocolate causes Nobel-level achievement. Country-level relationships cannot be transferred to individuals, and observational cross-country comparisons are vulnerable to confounding, measurement differences, and influential observations.
National income, education and research investment, population history, data quality, and other country-level factors are plausible alternative explanations, but this small ecological comparison does not identify which factor explains the association. Publication venue does not turn an ecological correlation into a causal estimate.
6. Autism Identification and Organic-Food Sales
A widely circulated teaching graph places two U.S. time series on the same chart: reported autism identification and organic-food sales. Both rise over part of the selected period. That visual alignment does not demonstrate a causal relationship, and this article found no credible evidence that organic-food consumption causes autism.
This is primarily a shared-trend problem. A responsible analysis would verify both source series, align definitions and units, disclose selected years, examine changes rather than levels, account for autocorrelation, and test a causal model before making any health claim.
CDC cautions that changes in measured autism prevalence reflect multiple factors, including diagnostic practices, screening, awareness, access to services, population coverage, and potentially other influences. Use precise language such as “identified prevalence” or “reported prevalence.”
Shoe size and reading ability is a standard classroom example, often demonstrated with toy or classroom data. Across a mixed-age sample of children, both measures can increase with age and schooling, creating an association that should not be interpreted as an effect of shoe size on reading.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This example demonstrates multiple-comparison and selection effects: searching many series makes extreme sample correlations more likely even when variables are unrelated. “P-hacking” is broader and should not be used as a synonym for every cherry-picked correlation. Report the search process, number of comparisons, uncertainty, and out-of-sample replication.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
The seven examples below are not all empirical studies. Some are sourced demonstrations, some are satire built from public series, and some are teaching hypotheticals. For each, ask who produced the data, how variables were defined, whether the window was selected in advance, how many comparisons were searched, what causal graph is plausible, and whether the relationship persists in new data.
1. Ice Cream Sales and Drowning Deaths

- Always consider seasonal factors when analyzing time-series data. Seasonal variations can create misleading correlations between variables.
- Look for third variables that might explain the observed correlation. Conduct thorough exploratory data analysis to identify potential confounding variables.
- Consider time-based confounding variables. Like seasonality, other time-related factors can create spurious relationships.
- Use this classic example to teach and reinforce the difference between correlation and causation. It's a simple yet powerful way to communicate this crucial statistical concept.
2. Divorce Rate in Maine vs Margarine Consumption

3. Number of Pirates vs Global Temperature
The pirates-versus-temperature graphic originated as satire associated with Bobby Henderson’s Flying Spaghetti Monster argument. It juxtaposes rough pirate-count estimates with temperature values to mock the leap from temporal association to causation. It is not a validated historical dataset or climate-attribution analysis.
Tyler Vigen published a deliberately selected annual series pairing Nicolas Cage film appearances with swimming-pool drownings. Treat the displayed coefficient and date range as properties of that selected sample. Extending the window, changing definitions, or testing new data can change the relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This example demonstrates multiple-comparison and selection effects: searching many series makes extreme sample correlations more likely even when variables are unrelated. “P-hacking” is broader and should not be used as a synonym for every cherry-picked correlation. Report the search process, number of comparisons, uncertainty, and out-of-sample replication.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
The seven examples below are not all empirical studies. Some are sourced demonstrations, some are satire built from public series, and some are teaching hypotheticals. For each, ask who produced the data, how variables were defined, whether the window was selected in advance, how many comparisons were searched, what causal graph is plausible, and whether the relationship persists in new data.
1. Ice Cream Sales and Drowning Deaths

- Always consider seasonal factors when analyzing time-series data. Seasonal variations can create misleading correlations between variables.
- Look for third variables that might explain the observed correlation. Conduct thorough exploratory data analysis to identify potential confounding variables.
- Consider time-based confounding variables. Like seasonality, other time-related factors can create spurious relationships.
- Use this classic example to teach and reinforce the difference between correlation and causation. It's a simple yet powerful way to communicate this crucial statistical concept.
2. Divorce Rate in Maine vs Margarine Consumption

3. Number of Pirates vs Global Temperature
The pirates-versus-temperature graphic originated as satire associated with Bobby Henderson’s Flying Spaghetti Monster argument. It juxtaposes rough pirate-count estimates with temperature values to mock the leap from temporal association to causation. It is not a validated historical dataset or climate-attribution analysis.
Use it only as a rhetorical illustration. Evidence about climate causes comes from physical measurements, attribution methods, and converging climate science—not from accepting or rejecting this cartoon correlation.
4. Nicolas Cage Movies and Swimming Pool Drownings

- Selected annual-series correlation: The close tracking of the two variables over a significant period provides a striking visual representation of a spurious correlation.
- Involves Entertainment Industry and Safety Statistics: The juxtaposition of these seemingly unrelated fields further highlights the randomness of the correlation.
- Demonstrates Coincidental Trending Patterns: This example perfectly illustrates how unrelated trends can align by chance.
- Easy to Visualize and Understand: The simplicity of the comparison makes it readily accessible to a broad audience, including those without a strong statistical background.
- Entertaining Way to Teach Statistical Concepts: The inherent humor of the example makes it an engaging way to introduce the concept of spurious correlations.
- Shows Absurdity of Assuming Causation from Correlation: This example effectively dismantles the common misconception that correlation implies causation.
- Memorable Celebrity Connection Aids Retention: The inclusion of Nicolas Cage adds a memorable element that strengthens the learning experience.
- May Trivialize Actual Drowning Statistics: Using a serious topic like accidental drownings in a humorous context could be perceived as insensitive.
- Represents Selective Data Mining: It highlights the potential for misleading results when data is selectively chosen to support a particular narrative.
- Could be Misunderstood as Intentionally Meaningful: Some individuals might misinterpret the correlation as having some underlying, albeit bizarre, meaning.
- Question correlations involving celebrities or pop culture: These are often ripe for spurious relationships due to the high visibility and frequent media coverage of these figures.
- Consider sample sizes and time periods in correlation analysis: Limited data sets and specific timeframes can create misleading correlations.
- Look for logical explanatory mechanisms: If a correlation lacks a plausible explanation, it should be treated with skepticism.
- Use entertaining examples like this to engage students while teaching serious concepts: This example provides a valuable lesson in critical thinking and data analysis.
5. Chocolate Consumption and Nobel Prize Winners
In a 2012 New England Journal of Medicine correspondence, Franz Messerli compared national chocolate consumption with Nobel laureates per capita and reported a country-level correlation. It did not establish that eating chocolate causes Nobel-level achievement. Country-level relationships cannot be transferred to individuals, and observational cross-country comparisons are vulnerable to confounding, measurement differences, and influential observations.
National income, education and research investment, population history, data quality, and other country-level factors are plausible alternative explanations, but this small ecological comparison does not identify which factor explains the association. Publication venue does not turn an ecological correlation into a causal estimate.
6. Autism Identification and Organic-Food Sales
A widely circulated teaching graph places two U.S. time series on the same chart: reported autism identification and organic-food sales. Both rise over part of the selected period. That visual alignment does not demonstrate a causal relationship, and this article found no credible evidence that organic-food consumption causes autism.
This is primarily a shared-trend problem. A responsible analysis would verify both source series, align definitions and units, disclose selected years, examine changes rather than levels, account for autocorrelation, and test a causal model before making any health claim.
CDC cautions that changes in measured autism prevalence reflect multiple factors, including diagnostic practices, screening, awareness, access to services, population coverage, and potentially other influences. Use precise language such as “identified prevalence” or “reported prevalence.”
Continue with Data4AI’s archive on statistical reasoning, data quality, and responsible model evaluation.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This example demonstrates multiple-comparison and selection effects: searching many series makes extreme sample correlations more likely even when variables are unrelated. “P-hacking” is broader and should not be used as a synonym for every cherry-picked correlation. Report the search process, number of comparisons, uncertainty, and out-of-sample replication.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
The seven examples below are not all empirical studies. Some are sourced demonstrations, some are satire built from public series, and some are teaching hypotheticals. For each, ask who produced the data, how variables were defined, whether the window was selected in advance, how many comparisons were searched, what causal graph is plausible, and whether the relationship persists in new data.
1. Ice Cream Sales and Drowning Deaths

- Always consider seasonal factors when analyzing time-series data. Seasonal variations can create misleading correlations between variables.
- Look for third variables that might explain the observed correlation. Conduct thorough exploratory data analysis to identify potential confounding variables.
- Consider time-based confounding variables. Like seasonality, other time-related factors can create spurious relationships.
- Use this classic example to teach and reinforce the difference between correlation and causation. It's a simple yet powerful way to communicate this crucial statistical concept.
2. Divorce Rate in Maine vs Margarine Consumption

3. Number of Pirates vs Global Temperature
The pirates-versus-temperature graphic originated as satire associated with Bobby Henderson’s Flying Spaghetti Monster argument. It juxtaposes rough pirate-count estimates with temperature values to mock the leap from temporal association to causation. It is not a validated historical dataset or climate-attribution analysis.
Tyler Vigen published a deliberately selected annual series pairing Nicolas Cage film appearances with swimming-pool drownings. Treat the displayed coefficient and date range as properties of that selected sample. Extending the window, changing definitions, or testing new data can change the relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This example demonstrates multiple-comparison and selection effects: searching many series makes extreme sample correlations more likely even when variables are unrelated. “P-hacking” is broader and should not be used as a synonym for every cherry-picked correlation. Report the search process, number of comparisons, uncertainty, and out-of-sample replication.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
The seven examples below are not all empirical studies. Some are sourced demonstrations, some are satire built from public series, and some are teaching hypotheticals. For each, ask who produced the data, how variables were defined, whether the window was selected in advance, how many comparisons were searched, what causal graph is plausible, and whether the relationship persists in new data.
1. Ice Cream Sales and Drowning Deaths

- Always consider seasonal factors when analyzing time-series data. Seasonal variations can create misleading correlations between variables.
- Look for third variables that might explain the observed correlation. Conduct thorough exploratory data analysis to identify potential confounding variables.
- Consider time-based confounding variables. Like seasonality, other time-related factors can create spurious relationships.
- Use this classic example to teach and reinforce the difference between correlation and causation. It's a simple yet powerful way to communicate this crucial statistical concept.
2. Divorce Rate in Maine vs Margarine Consumption

3. Number of Pirates vs Global Temperature
The pirates-versus-temperature graphic originated as satire associated with Bobby Henderson’s Flying Spaghetti Monster argument. It juxtaposes rough pirate-count estimates with temperature values to mock the leap from temporal association to causation. It is not a validated historical dataset or climate-attribution analysis.
Use it only as a rhetorical illustration. Evidence about climate causes comes from physical measurements, attribution methods, and converging climate science—not from accepting or rejecting this cartoon correlation.
4. Nicolas Cage Movies and Swimming Pool Drownings

- Selected annual-series correlation: The close tracking of the two variables over a significant period provides a striking visual representation of a spurious correlation.
- Involves Entertainment Industry and Safety Statistics: The juxtaposition of these seemingly unrelated fields further highlights the randomness of the correlation.
- Demonstrates Coincidental Trending Patterns: This example perfectly illustrates how unrelated trends can align by chance.
- Easy to Visualize and Understand: The simplicity of the comparison makes it readily accessible to a broad audience, including those without a strong statistical background.
- Entertaining Way to Teach Statistical Concepts: The inherent humor of the example makes it an engaging way to introduce the concept of spurious correlations.
- Shows Absurdity of Assuming Causation from Correlation: This example effectively dismantles the common misconception that correlation implies causation.
- Memorable Celebrity Connection Aids Retention: The inclusion of Nicolas Cage adds a memorable element that strengthens the learning experience.
- May Trivialize Actual Drowning Statistics: Using a serious topic like accidental drownings in a humorous context could be perceived as insensitive.
- Represents Selective Data Mining: It highlights the potential for misleading results when data is selectively chosen to support a particular narrative.
- Could be Misunderstood as Intentionally Meaningful: Some individuals might misinterpret the correlation as having some underlying, albeit bizarre, meaning.
- Question correlations involving celebrities or pop culture: These are often ripe for spurious relationships due to the high visibility and frequent media coverage of these figures.
- Consider sample sizes and time periods in correlation analysis: Limited data sets and specific timeframes can create misleading correlations.
- Look for logical explanatory mechanisms: If a correlation lacks a plausible explanation, it should be treated with skepticism.
- Use entertaining examples like this to engage students while teaching serious concepts: This example provides a valuable lesson in critical thinking and data analysis.
5. Chocolate Consumption and Nobel Prize Winners
In a 2012 New England Journal of Medicine correspondence, Franz Messerli compared national chocolate consumption with Nobel laureates per capita and reported a country-level correlation. It did not establish that eating chocolate causes Nobel-level achievement. Country-level relationships cannot be transferred to individuals, and observational cross-country comparisons are vulnerable to confounding, measurement differences, and influential observations.
National income, education and research investment, population history, data quality, and other country-level factors are plausible alternative explanations, but this small ecological comparison does not identify which factor explains the association. Publication venue does not turn an ecological correlation into a causal estimate.
6. Autism Identification and Organic-Food Sales
A widely circulated teaching graph places two U.S. time series on the same chart: reported autism identification and organic-food sales. Both rise over part of the selected period. That visual alignment does not demonstrate a causal relationship, and this article found no credible evidence that organic-food consumption causes autism.
This is primarily a shared-trend problem. A responsible analysis would verify both source series, align definitions and units, disclose selected years, examine changes rather than levels, account for autocorrelation, and test a causal model before making any health claim.
CDC cautions that changes in measured autism prevalence reflect multiple factors, including diagnostic practices, screening, awareness, access to services, population coverage, and potentially other influences. Use precise language such as “identified prevalence” or “reported prevalence.”
Shoe size and reading ability is a standard classroom example, often demonstrated with toy or classroom data. Across a mixed-age sample of children, both measures can increase with age and schooling, creating an association that should not be interpreted as an effect of shoe size on reading.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This example demonstrates multiple-comparison and selection effects: searching many series makes extreme sample correlations more likely even when variables are unrelated. “P-hacking” is broader and should not be used as a synonym for every cherry-picked correlation. Report the search process, number of comparisons, uncertainty, and out-of-sample replication.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
The seven examples below are not all empirical studies. Some are sourced demonstrations, some are satire built from public series, and some are teaching hypotheticals. For each, ask who produced the data, how variables were defined, whether the window was selected in advance, how many comparisons were searched, what causal graph is plausible, and whether the relationship persists in new data.
1. Ice Cream Sales and Drowning Deaths

- Always consider seasonal factors when analyzing time-series data. Seasonal variations can create misleading correlations between variables.
- Look for third variables that might explain the observed correlation. Conduct thorough exploratory data analysis to identify potential confounding variables.
- Consider time-based confounding variables. Like seasonality, other time-related factors can create spurious relationships.
- Use this classic example to teach and reinforce the difference between correlation and causation. It's a simple yet powerful way to communicate this crucial statistical concept.
2. Divorce Rate in Maine vs Margarine Consumption

3. Number of Pirates vs Global Temperature
The pirates-versus-temperature graphic originated as satire associated with Bobby Henderson’s Flying Spaghetti Monster argument. It juxtaposes rough pirate-count estimates with temperature values to mock the leap from temporal association to causation. It is not a validated historical dataset or climate-attribution analysis.
Tyler Vigen published a deliberately selected annual series pairing Nicolas Cage film appearances with swimming-pool drownings. Treat the displayed coefficient and date range as properties of that selected sample. Extending the window, changing definitions, or testing new data can change the relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This example demonstrates multiple-comparison and selection effects: searching many series makes extreme sample correlations more likely even when variables are unrelated. “P-hacking” is broader and should not be used as a synonym for every cherry-picked correlation. Report the search process, number of comparisons, uncertainty, and out-of-sample replication.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
The seven examples below are not all empirical studies. Some are sourced demonstrations, some are satire built from public series, and some are teaching hypotheticals. For each, ask who produced the data, how variables were defined, whether the window was selected in advance, how many comparisons were searched, what causal graph is plausible, and whether the relationship persists in new data.
1. Ice Cream Sales and Drowning Deaths

- Always consider seasonal factors when analyzing time-series data. Seasonal variations can create misleading correlations between variables.
- Look for third variables that might explain the observed correlation. Conduct thorough exploratory data analysis to identify potential confounding variables.
- Consider time-based confounding variables. Like seasonality, other time-related factors can create spurious relationships.
- Use this classic example to teach and reinforce the difference between correlation and causation. It's a simple yet powerful way to communicate this crucial statistical concept.
2. Divorce Rate in Maine vs Margarine Consumption

3. Number of Pirates vs Global Temperature
The pirates-versus-temperature graphic originated as satire associated with Bobby Henderson’s Flying Spaghetti Monster argument. It juxtaposes rough pirate-count estimates with temperature values to mock the leap from temporal association to causation. It is not a validated historical dataset or climate-attribution analysis.
Use it only as a rhetorical illustration. Evidence about climate causes comes from physical measurements, attribution methods, and converging climate science—not from accepting or rejecting this cartoon correlation.
4. Nicolas Cage Movies and Swimming Pool Drownings

- Selected annual-series correlation: The close tracking of the two variables over a significant period provides a striking visual representation of a spurious correlation.
- Involves Entertainment Industry and Safety Statistics: The juxtaposition of these seemingly unrelated fields further highlights the randomness of the correlation.
- Demonstrates Coincidental Trending Patterns: This example perfectly illustrates how unrelated trends can align by chance.
- Easy to Visualize and Understand: The simplicity of the comparison makes it readily accessible to a broad audience, including those without a strong statistical background.
- Entertaining Way to Teach Statistical Concepts: The inherent humor of the example makes it an engaging way to introduce the concept of spurious correlations.
- Shows Absurdity of Assuming Causation from Correlation: This example effectively dismantles the common misconception that correlation implies causation.
- Memorable Celebrity Connection Aids Retention: The inclusion of Nicolas Cage adds a memorable element that strengthens the learning experience.
- May Trivialize Actual Drowning Statistics: Using a serious topic like accidental drownings in a humorous context could be perceived as insensitive.
- Represents Selective Data Mining: It highlights the potential for misleading results when data is selectively chosen to support a particular narrative.
- Could be Misunderstood as Intentionally Meaningful: Some individuals might misinterpret the correlation as having some underlying, albeit bizarre, meaning.
- Question correlations involving celebrities or pop culture: These are often ripe for spurious relationships due to the high visibility and frequent media coverage of these figures.
- Consider sample sizes and time periods in correlation analysis: Limited data sets and specific timeframes can create misleading correlations.
- Look for logical explanatory mechanisms: If a correlation lacks a plausible explanation, it should be treated with skepticism.
- Use entertaining examples like this to engage students while teaching serious concepts: This example provides a valuable lesson in critical thinking and data analysis.
5. Chocolate Consumption and Nobel Prize Winners
In a 2012 New England Journal of Medicine correspondence, Franz Messerli compared national chocolate consumption with Nobel laureates per capita and reported a country-level correlation. It did not establish that eating chocolate causes Nobel-level achievement. Country-level relationships cannot be transferred to individuals, and observational cross-country comparisons are vulnerable to confounding, measurement differences, and influential observations.
National income, education and research investment, population history, data quality, and other country-level factors are plausible alternative explanations, but this small ecological comparison does not identify which factor explains the association. Publication venue does not turn an ecological correlation into a causal estimate.
6. Autism Identification and Organic-Food Sales
A widely circulated teaching graph places two U.S. time series on the same chart: reported autism identification and organic-food sales. Both rise over part of the selected period. That visual alignment does not demonstrate a causal relationship, and this article found no credible evidence that organic-food consumption causes autism.
This is primarily a shared-trend problem. A responsible analysis would verify both source series, align definitions and units, disclose selected years, examine changes rather than levels, account for autocorrelation, and test a causal model before making any health claim.
CDC cautions that changes in measured autism prevalence reflect multiple factors, including diagnostic practices, screening, awareness, access to services, population coverage, and potentially other influences. Use precise language such as “identified prevalence” or “reported prevalence.”
Adjusting or stratifying by age and educational stage can reduce or remove the association in the teaching example, but that result must be shown in the actual dataset rather than asserted universally. Conditioning should be justified by a causal model; mechanically controlling for everything can introduce bias.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This example demonstrates multiple-comparison and selection effects: searching many series makes extreme sample correlations more likely even when variables are unrelated. “P-hacking” is broader and should not be used as a synonym for every cherry-picked correlation. Report the search process, number of comparisons, uncertainty, and out-of-sample replication.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
The seven examples below are not all empirical studies. Some are sourced demonstrations, some are satire built from public series, and some are teaching hypotheticals. For each, ask who produced the data, how variables were defined, whether the window was selected in advance, how many comparisons were searched, what causal graph is plausible, and whether the relationship persists in new data.
1. Ice Cream Sales and Drowning Deaths

- Always consider seasonal factors when analyzing time-series data. Seasonal variations can create misleading correlations between variables.
- Look for third variables that might explain the observed correlation. Conduct thorough exploratory data analysis to identify potential confounding variables.
- Consider time-based confounding variables. Like seasonality, other time-related factors can create spurious relationships.
- Use this classic example to teach and reinforce the difference between correlation and causation. It's a simple yet powerful way to communicate this crucial statistical concept.
2. Divorce Rate in Maine vs Margarine Consumption

3. Number of Pirates vs Global Temperature
The pirates-versus-temperature graphic originated as satire associated with Bobby Henderson’s Flying Spaghetti Monster argument. It juxtaposes rough pirate-count estimates with temperature values to mock the leap from temporal association to causation. It is not a validated historical dataset or climate-attribution analysis.
Tyler Vigen published a deliberately selected annual series pairing Nicolas Cage film appearances with swimming-pool drownings. Treat the displayed coefficient and date range as properties of that selected sample. Extending the window, changing definitions, or testing new data can change the relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This example demonstrates multiple-comparison and selection effects: searching many series makes extreme sample correlations more likely even when variables are unrelated. “P-hacking” is broader and should not be used as a synonym for every cherry-picked correlation. Report the search process, number of comparisons, uncertainty, and out-of-sample replication.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
The seven examples below are not all empirical studies. Some are sourced demonstrations, some are satire built from public series, and some are teaching hypotheticals. For each, ask who produced the data, how variables were defined, whether the window was selected in advance, how many comparisons were searched, what causal graph is plausible, and whether the relationship persists in new data.
1. Ice Cream Sales and Drowning Deaths

- Always consider seasonal factors when analyzing time-series data. Seasonal variations can create misleading correlations between variables.
- Look for third variables that might explain the observed correlation. Conduct thorough exploratory data analysis to identify potential confounding variables.
- Consider time-based confounding variables. Like seasonality, other time-related factors can create spurious relationships.
- Use this classic example to teach and reinforce the difference between correlation and causation. It's a simple yet powerful way to communicate this crucial statistical concept.
2. Divorce Rate in Maine vs Margarine Consumption

3. Number of Pirates vs Global Temperature
The pirates-versus-temperature graphic originated as satire associated with Bobby Henderson’s Flying Spaghetti Monster argument. It juxtaposes rough pirate-count estimates with temperature values to mock the leap from temporal association to causation. It is not a validated historical dataset or climate-attribution analysis.
Use it only as a rhetorical illustration. Evidence about climate causes comes from physical measurements, attribution methods, and converging climate science—not from accepting or rejecting this cartoon correlation.
4. Nicolas Cage Movies and Swimming Pool Drownings

- Selected annual-series correlation: The close tracking of the two variables over a significant period provides a striking visual representation of a spurious correlation.
- Involves Entertainment Industry and Safety Statistics: The juxtaposition of these seemingly unrelated fields further highlights the randomness of the correlation.
- Demonstrates Coincidental Trending Patterns: This example perfectly illustrates how unrelated trends can align by chance.
- Easy to Visualize and Understand: The simplicity of the comparison makes it readily accessible to a broad audience, including those without a strong statistical background.
- Entertaining Way to Teach Statistical Concepts: The inherent humor of the example makes it an engaging way to introduce the concept of spurious correlations.
- Shows Absurdity of Assuming Causation from Correlation: This example effectively dismantles the common misconception that correlation implies causation.
- Memorable Celebrity Connection Aids Retention: The inclusion of Nicolas Cage adds a memorable element that strengthens the learning experience.
- May Trivialize Actual Drowning Statistics: Using a serious topic like accidental drownings in a humorous context could be perceived as insensitive.
- Represents Selective Data Mining: It highlights the potential for misleading results when data is selectively chosen to support a particular narrative.
- Could be Misunderstood as Intentionally Meaningful: Some individuals might misinterpret the correlation as having some underlying, albeit bizarre, meaning.
- Question correlations involving celebrities or pop culture: These are often ripe for spurious relationships due to the high visibility and frequent media coverage of these figures.
- Consider sample sizes and time periods in correlation analysis: Limited data sets and specific timeframes can create misleading correlations.
- Look for logical explanatory mechanisms: If a correlation lacks a plausible explanation, it should be treated with skepticism.
- Use entertaining examples like this to engage students while teaching serious concepts: This example provides a valuable lesson in critical thinking and data analysis.
5. Chocolate Consumption and Nobel Prize Winners
In a 2012 New England Journal of Medicine correspondence, Franz Messerli compared national chocolate consumption with Nobel laureates per capita and reported a country-level correlation. It did not establish that eating chocolate causes Nobel-level achievement. Country-level relationships cannot be transferred to individuals, and observational cross-country comparisons are vulnerable to confounding, measurement differences, and influential observations.
National income, education and research investment, population history, data quality, and other country-level factors are plausible alternative explanations, but this small ecological comparison does not identify which factor explains the association. Publication venue does not turn an ecological correlation into a causal estimate.
6. Autism Identification and Organic-Food Sales
A widely circulated teaching graph places two U.S. time series on the same chart: reported autism identification and organic-food sales. Both rise over part of the selected period. That visual alignment does not demonstrate a causal relationship, and this article found no credible evidence that organic-food consumption causes autism.
This is primarily a shared-trend problem. A responsible analysis would verify both source series, align definitions and units, disclose selected years, examine changes rather than levels, account for autocorrelation, and test a causal model before making any health claim.
CDC cautions that changes in measured autism prevalence reflect multiple factors, including diagnostic practices, screening, awareness, access to services, population coverage, and potentially other influences. Use precise language such as “identified prevalence” or “reported prevalence.”
Shoe size and reading ability is a standard classroom example, often demonstrated with toy or classroom data. Across a mixed-age sample of children, both measures can increase with age and schooling, creating an association that should not be interpreted as an effect of shoe size on reading.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This example demonstrates multiple-comparison and selection effects: searching many series makes extreme sample correlations more likely even when variables are unrelated. “P-hacking” is broader and should not be used as a synonym for every cherry-picked correlation. Report the search process, number of comparisons, uncertainty, and out-of-sample replication.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
The seven examples below are not all empirical studies. Some are sourced demonstrations, some are satire built from public series, and some are teaching hypotheticals. For each, ask who produced the data, how variables were defined, whether the window was selected in advance, how many comparisons were searched, what causal graph is plausible, and whether the relationship persists in new data.
1. Ice Cream Sales and Drowning Deaths

- Always consider seasonal factors when analyzing time-series data. Seasonal variations can create misleading correlations between variables.
- Look for third variables that might explain the observed correlation. Conduct thorough exploratory data analysis to identify potential confounding variables.
- Consider time-based confounding variables. Like seasonality, other time-related factors can create spurious relationships.
- Use this classic example to teach and reinforce the difference between correlation and causation. It's a simple yet powerful way to communicate this crucial statistical concept.
2. Divorce Rate in Maine vs Margarine Consumption

3. Number of Pirates vs Global Temperature
The pirates-versus-temperature graphic originated as satire associated with Bobby Henderson’s Flying Spaghetti Monster argument. It juxtaposes rough pirate-count estimates with temperature values to mock the leap from temporal association to causation. It is not a validated historical dataset or climate-attribution analysis.
Tyler Vigen published a deliberately selected annual series pairing Nicolas Cage film appearances with swimming-pool drownings. Treat the displayed coefficient and date range as properties of that selected sample. Extending the window, changing definitions, or testing new data can change the relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This example demonstrates multiple-comparison and selection effects: searching many series makes extreme sample correlations more likely even when variables are unrelated. “P-hacking” is broader and should not be used as a synonym for every cherry-picked correlation. Report the search process, number of comparisons, uncertainty, and out-of-sample replication.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
Tyler Vigen’s selected 2000–2009 series reports a Pearson correlation of r = 0.9925585 between U.S. per-capita margarine consumption and Maine’s divorce rate. That is a descriptive result for ten paired annual observations selected from a very large search; it is not evidence that either variable causes the other.
A correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
This is a common teaching example rather than evidence that a single published dataset always produces the same coefficient. Season, temperature, and exposure to swimming are plausible common causes. Actual drowning risk also depends on supervision, swimming ability, alcohol, location, safety measures, and other factors. The example illustrates a hypothesis about confounding; it does not establish one universal relationship.
>Why a Correlation Can MisleadA correlation summarizes how two measured variables vary together; by itself it does not identify why they move together. A noncausal association can arise from a common cause, selection bias, reverse direction, shared time trends, measurement choices, or chance—especially when analysts search many variable pairs and report only striking results.
The seven examples below are not all empirical studies. Some are sourced demonstrations, some are satire built from public series, and some are teaching hypotheticals. For each, ask who produced the data, how variables were defined, whether the window was selected in advance, how many comparisons were searched, what causal graph is plausible, and whether the relationship persists in new data.
1. Ice Cream Sales and Drowning Deaths

- Always consider seasonal factors when analyzing time-series data. Seasonal variations can create misleading correlations between variables.
- Look for third variables that might explain the observed correlation. Conduct thorough exploratory data analysis to identify potential confounding variables.
- Consider time-based confounding variables. Like seasonality, other time-related factors can create spurious relationships.
- Use this classic example to teach and reinforce the difference between correlation and causation. It's a simple yet powerful way to communicate this crucial statistical concept.
2. Divorce Rate in Maine vs Margarine Consumption

3. Number of Pirates vs Global Temperature
The pirates-versus-temperature graphic originated as satire associated with Bobby Henderson’s Flying Spaghetti Monster argument. It juxtaposes rough pirate-count estimates with temperature values to mock the leap from temporal association to causation. It is not a validated historical dataset or climate-attribution analysis.
Use it only as a rhetorical illustration. Evidence about climate causes comes from physical measurements, attribution methods, and converging climate science—not from accepting or rejecting this cartoon correlation.
4. Nicolas Cage Movies and Swimming Pool Drownings

- Selected annual-series correlation: The close tracking of the two variables over a significant period provides a striking visual representation of a spurious correlation.
- Involves Entertainment Industry and Safety Statistics: The juxtaposition of these seemingly unrelated fields further highlights the randomness of the correlation.
- Demonstrates Coincidental Trending Patterns: This example perfectly illustrates how unrelated trends can align by chance.
- Easy to Visualize and Understand: The simplicity of the comparison makes it readily accessible to a broad audience, including those without a strong statistical background.
- Entertaining Way to Teach Statistical Concepts: The inherent humor of the example makes it an engaging way to introduce the concept of spurious correlations.
- Shows Absurdity of Assuming Causation from Correlation: This example effectively dismantles the common misconception that correlation implies causation.
- Memorable Celebrity Connection Aids Retention: The inclusion of Nicolas Cage adds a memorable element that strengthens the learning experience.
- May Trivialize Actual Drowning Statistics: Using a serious topic like accidental drownings in a humorous context could be perceived as insensitive.
- Represents Selective Data Mining: It highlights the potential for misleading results when data is selectively chosen to support a particular narrative.
- Could be Misunderstood as Intentionally Meaningful: Some individuals might misinterpret the correlation as having some underlying, albeit bizarre, meaning.
- Question correlations involving celebrities or pop culture: These are often ripe for spurious relationships due to the high visibility and frequent media coverage of these figures.
- Consider sample sizes and time periods in correlation analysis: Limited data sets and specific timeframes can create misleading correlations.
- Look for logical explanatory mechanisms: If a correlation lacks a plausible explanation, it should be treated with skepticism.
- Use entertaining examples like this to engage students while teaching serious concepts: This example provides a valuable lesson in critical thinking and data analysis.
5. Chocolate Consumption and Nobel Prize Winners
In a 2012 New England Journal of Medicine correspondence, Franz Messerli compared national chocolate consumption with Nobel laureates per capita and reported a country-level correlation. It did not establish that eating chocolate causes Nobel-level achievement. Country-level relationships cannot be transferred to individuals, and observational cross-country comparisons are vulnerable to confounding, measurement differences, and influential observations.
National income, education and research investment, population history, data quality, and other country-level factors are plausible alternative explanations, but this small ecological comparison does not identify which factor explains the association. Publication venue does not turn an ecological correlation into a causal estimate.
6. Autism Identification and Organic-Food Sales
A widely circulated teaching graph places two U.S. time series on the same chart: reported autism identification and organic-food sales. Both rise over part of the selected period. That visual alignment does not demonstrate a causal relationship, and this article found no credible evidence that organic-food consumption causes autism.
This is primarily a shared-trend problem. A responsible analysis would verify both source series, align definitions and units, disclose selected years, examine changes rather than levels, account for autocorrelation, and test a causal model before making any health claim.
CDC cautions that changes in measured autism prevalence reflect multiple factors, including diagnostic practices, screening, awareness, access to services, population coverage, and potentially other influences. Use precise language such as “identified prevalence” or “reported prevalence.”
Evidence classification: teaching illustration; not a causal health study.
7. Shoe Size and Reading Ability in Children
This classic example of spurious correlation serves as a crucial lesson for anyone working with data, especially in fields like AI research, machine learning, and data science. It highlights the critical importance of understanding the underlying relationships between variables and the dangers of drawing conclusions based solely on observed correlations. While seemingly trivial, the relationship between shoe size and reading ability in children perfectly encapsulates the concept of a confounding variable and provides a clear, easy-to-understand illustration of why controlling for such variables is essential for accurate analysis. This makes it a valuable example for data scientists, AI researchers, and anyone interpreting statistical data. The observed correlation is this: studies consistently show that children with larger shoe sizes tend to perform better on reading tests. On the surface, this might lead one to entertain outlandish theories about foot growth somehow influencing cognitive development. However, the relationship vanishes when we introduce a third variable: age. Both shoe size and reading ability naturally increase with age. Older children, having had more time for both physical and cognitive development, tend to have larger feet and stronger reading skills than younger children. Therefore, the initial correlation between shoe size and reading ability isn't a causal relationship but an artifact of their shared association with age – a classic example of a spurious correlation. This spurious correlation appears consistently across various educational studies, making it a robust example for demonstrating the concept. It's particularly relevant because it connects physical development (shoe size) with a cognitive skill (reading ability), showcasing how easily seemingly unrelated factors can appear linked. This example’s strength lies in its simplicity; it’s easy to visualize and grasp, even for those without a strong statistical background. This makes it an excellent pedagogical tool in educational research methodology courses, psychology statistics textbooks, and teacher training programs. The ease with which this spurious correlation can be controlled further enhances its educational value. By simply including age as a control variable in the analysis, the illusory relationship between shoe size and reading ability disappears. This demonstrates, in a practical way, how statistical controls can help uncover the true nature of relationships between variables. It highlights the importance of considering developmental factors, particularly when working with data related to children, and underscores the need to account for variables that naturally increase together over time. While a powerful teaching tool, the shoe size and reading ability example does have some limitations. For experienced data scientists or AI researchers, the example might seem overly simplistic or even obvious. The focus on this correlation could also inadvertently oversimplify the complex reality of educational assessment, potentially distracting from other genuine factors influencing reading ability, such as socioeconomic background, access to resources, and individual learning differences. Actionable Tips for Data Professionals:- Always consider age and developmental stage in child-related studies. This is crucial for avoiding misinterpretations of correlational data.
- Use statistical controls to test for confounding variables. Techniques like regression analysis allow you to isolate the effect of one variable while controlling for the influence of others.
- Look for variables that naturally increase together over time. These are prime candidates for creating spurious correlations if not properly controlled.
- Document and explain your methodology clearly, especially when dealing with potential confounding variables. This ensures transparency and allows others to replicate and validate your findings.
- When building AI models, be cautious about including features that might lead to spurious correlations. Thoroughly analyze your data and consider the underlying relationships between variables before incorporating them into your model.
7 Notable Spurious Correlation Examples Comparison
| Example | Evidence type | Main statistical trap | What to check |
|---|---|---|---|
| Ice cream and drowning | Teaching example | Seasonal/common-cause confounding | Season, weather, exposure, safety |
| Margarine and Maine divorce | Selected public time series | Mass comparison and selection | Search universe, window, replication |
| Pirates and temperature | Satire | Shared trend and implausible mechanism | Data provenance |
| Nicolas Cage and drownings | Selected public time series | Data dredging and window selection | Definitions and out-of-sample behavior |
| Chocolate and Nobel laureates | Published ecological correlation | Ecological inference and confounding | Country covariates and influential observations |
| Autism identification and organic sales | Teaching graph | Correlating trending series | Definitions, detrending, autocorrelation |
| Shoe size and reading | Teaching example | Age and schooling | Causally justified adjustment |

Historical comments from Datanizant
No public comments on this article
No approved public comments were included in the WordPress export for this article.