Correlation, Causation
| Stage | Tool |
|---|---|
| Knowledge | Types of knowledge; what makes knowledge scientific |
| Arguments | Premises, conclusions, argument structure |
| Logic | Validity, soundness, deduction vs. induction |
| Failure modes | Fallacies, misinformation, persuasion |
Today we move from arguments to evidence: what happens when the premises are empirical claims about the world?
Why causation is always an inference
Drop a piece of chalk. What do you see?
Hume (1748): we never observe causation itself — only constant conjunction: A happens, then B happens, again and again.
“All reasonings concerning matter of fact seem to be founded on the relation of Cause and Effect… causes and effects are discoverable, not by reason but by experience.”
— Hume, Enquiry, Section IV
Every causal claim — “smoking causes cancer,” “austerity caused the recession,” “social media causes polarization” — is an inference from observed patterns, not an observation.
Consequence: causal claims can always be wrong. The pattern is real; the story we tell about it might not be.
Four stories behind every pattern
Two variables are correlated when they move together:
A correlation is a fact about the data.
X is correlated with Y. Why?
Test: ask which way the arrow could plausibly run. If both directions tell a coherent story, the correlation alone cannot decide.
Ice cream sales and drownings rise together. Does ice cream cause drowning?
Summer causes both: heat drives ice cream sales and swimming.
Classic confounders in social science:
Test: ask “what third factor could produce both?”
With enough variables, absurd correlations are guaranteed:
If you test 100 variable pairs, roughly 5 will look “statistically significant” by pure luck.
This is not a joke about bad researchers — it is a mathematical property of searching for patterns. Remember it when we reach p-hacking.
A pattern in groups need not hold for the individuals in them:
A tool for testing every “therefore”
Arguments about evidence hide their weakest step behind inferential words:
therefore — thus — this shows — this proves — this rules out — this is inconsistent with — this confirms
These words assert that the evidence supports the conclusion. They do not demonstrate it.
The Inference Audit: every time you meet one of these words, stop and test the logic. Six questions, one rule.
Most bad empirical arguments share one structure:
The rule: Do not ask only whether the evidence fits the author’s story. Ask whether it also fits the story the author is trying to reject.
A passage from a realistic policy report:
“Some blame federalism for the slow economic convergence of poor regions. But this cannot be right: convergence was fastest under the earlier centralized system and slowed down after decentralization. Therefore, federalism does not explain slow convergence.”
Sounds rigorous. Run the audit.
The evidence the author uses to refute the federalism hypothesis is precisely what the federalism hypothesis predicts. The “therefore” is empty.
Name the mistake: failure to derive the rival hypothesis’s prediction.
This is confirmation bias, presented as data analysis.
Before accepting a substantive story, check whether the pattern could arise from arithmetic alone:
| Pattern | Possible mechanical cause |
|---|---|
| “Improvement slowed after the reform” | Ceiling effect, diminishing returns |
| “The worst performers improved the most” | Mean reversion |
| “Program participants did better” | Selection: who joins? |
| “Average scores fell as enrollment grew” | Composition change |
| “X rose over the decade, and so did Y” | Both follow a general time trend |
None of these requires the author’s story to be true — or false. They require the author to rule them out.
For every “therefore” in an argument, ask:
Final test: if the rival hypothesis predicts the same pattern we observe, the evidence cannot be used to reject that rival.
Sample, method, bias
A scientific study is not a fact — it is an argument: premises (data, methods) supporting a conclusion (findings).
So everything the course has taught about arguments — plus the Inference Audit — applies. Plus three study-specific questions:
Ask: who is in the sample, who is missing, and does the conclusion quietly extend beyond the people actually studied?
Bias does not mean fraud. It means the filter between all studies conducted and the studies you get to read is not neutral.
Recall: test 100 comparisons, and ~5 look significant by chance.
Now imagine a researcher who tries many outcomes, subgroups, and model specifications — and reports the one that “worked.”
This is the chance explanation from Part 2, at scale.
For each claim: identify the pattern, list the rival explanations (reverse causation, confounder, chance, mechanical).
Popescu (TEC): Logic and Scientific Thinking