Correlation Isn't Causation: What the Examples Actually Show
Two things can rise and fall together without either one causing the other. The phrase 'correlation does not imply causation' is true, but on its own it does not tell you what to check instead. What actually separates a genuine cause from a coincidence is a specific set of questions about timing, mechanism, and hidden third factors.
Why the slogan falls short
The phrase gets used as a conversation-stopper. Someone presents a pattern, someone else says 'correlation isn't causation', and the discussion ends there. That is unsatisfying, because the pattern might still be causal. The slogan tells you not to assume causation, but it does not tell you how to find out whether it is there.
Correlation simply means two things tend to move together. When one goes up, the other tends to go up or down too. Causation means one thing actually produces the change in the other. Mixing these up is easy because a real cause always produces a correlation, but a correlation can appear for several other reasons entirely.
Classic examples explained
Ice cream sales and drowning deaths rise and fall together across the year. Nobody thinks ice cream causes drowning. Both are driven by a third factor, hot weather, which brings more swimmers to the water and more customers to ice cream vans. This is the textbook case of a common cause creating a correlation between two unrelated things.
Children with larger shoe sizes tend to read better on average. Shoe size does not improve reading. Age is doing the work here: older children have bigger feet and more years of schooling. Strip out age and the correlation vanishes, which is a useful test in itself.
There are well-known joke datasets showing things like the number of films a particular actor appeared in tracking the number of people who drowned in swimming pools that year. These are flukes. With enough variables and enough years of data, some will line up by chance alone, with no connection of any kind.
Three ways correlations mislead
The first trap is reverse causation. Poor sleep and anxiety are correlated, but which causes which. Anxiety can keep people awake, and bad sleep can also worsen anxiety the next day. The correlation is real, the arrow of cause is the open question, and getting the direction wrong leads to the wrong fix.
The second trap is confounding, where a hidden third factor drives both variables, as with ice cream and drowning. Confounders are often less obvious than weather. Income, age, education and location quietly shape many health and social statistics, so a correlation between two things can really be a correlation between each of them and something else entirely.
The third trap is coincidence. When researchers test many variables against many other variables, some pairs will correlate strongly purely by chance. This is more common now that large datasets make it cheap to test thousands of combinations. A striking pattern found this way often disappears when checked against new data.
What causation actually requires
Showing a cause usually means satisfying several conditions together, not just one. The cause has to come before the effect in time. There needs to be a plausible mechanism, some story for how one thing could bring about the other. The pattern should hold up across different groups, places and time periods, not just in one dataset.
The strongest evidence comes from experiments where researchers change one thing deliberately and watch what happens, while everything else is kept as similar as possible. In a randomised trial, people are assigned by chance to get a treatment or not, which spreads confounders evenly between the groups. That is what makes the comparison fair.
When an experiment is not possible, for instance in studying the effects of smoking or poverty, researchers rely on statistical methods that try to adjust for known confounders, compare natural experiments, or track the same people over time. None of these give the same certainty as a controlled trial, but they can build a strong case when several independent approaches point the same way.
A practical checklist
Before treating a correlation as a cause, it helps to ask a short list of questions. Could the arrow run the other way, with the supposed effect actually producing the supposed cause. Is there an obvious third factor, such as age, wealth or weather, that could explain both at once.
Does the pattern show up again in different data, different countries or different years, or does it only appear in one study. Is there a sensible mechanism connecting the two things, something that could be explained to a curious child without hand-waving. Was the data gathered by watching what naturally happens, or by an experiment that actually changed something.
None of these questions alone proves causation. Together, they shift a claim from 'these two things happen together' towards 'there is good reason to think one produces the other'. That shift, not the slogan itself, is the actual work of telling correlation from causation.
Common questions
- Can something be correlated and still be causal?
Yes, in fact every real causal relationship will also show up as a correlation. The problem is only that correlation alone cannot tell you whether a relationship is causal, coincidental, or driven by something else.
- What exactly is a confounding variable?
A confounding variable is a hidden factor that influences both things you are comparing, making them appear linked when they are not directly connected. Age, income and weather are common confounders in everyday statistics.
- Is a randomised controlled trial the only way to prove causation?
It is the strongest single method because random assignment spreads out confounders evenly, but it is not the only route. Repeated observational studies that control for known confounders, combined with a clear mechanism, can also build a convincing causal case.
- Why do news reports so often treat correlation as causation?
A single clean cause-and-effect story is easier to write and more interesting to read than a careful discussion of confounders and reverse causation. Pressure to publish quickly also means fewer checks get made before a claim is reported.
- How can I tell if a correlation I've seen online is a coincidence?
Look for whether the pattern has been tested again with new or different data, and whether anyone has proposed a believable mechanism connecting the two things. If it only appears in one dataset with no obvious explanation, treat it as likely chance.
Lucent, once a week
We read the new research and write up what actually holds, in the same plain language as this page. No jargon, no hype, three papers a week.
Read the latest issueRead next
- How to Tell If a Study Is Reliable in Five Minutes
A reliable study has a clear method, a fair comparison group, a sensible sample size, and no hidden conflicts of interest influencing the result.
- Preprints vs Peer Review: What Actually Changes
A preprint is a paper shared online before scientists check it, while a peer-reviewed paper has been scrutinised, challenged and usually revised first.
- A P-Value Explained Simply, Plus Three Common Mistakes
A p-value shows how likely your data would be if there were no real effect, not whether the finding is true, important, or unlikely to be chance.