Logic and Scientific Thinking

Correlation and Causation

Bogdan G. Popescu

Tecnologico de Monterrey

Welcome

Learning Outcomes

Overview

  1. Say what a causal claim actually asserts — and why it is never directly observed.
  1. Derive what each rival explanation predicts — and name the comparison that would discriminate between them.
  1. Read a scientific study critically: who was studied, what was measured, compared to what, and how the result was selected.

Roadmap

Where we are going today

  1. Hume’s problem: why causation is never observed — and what a causal claim really asserts
  2. Correlation vs. causation: the four default suspects behind any pattern
  3. The Inference Audit: testing every “therefore” — plus five patterns that need no cause at all
  4. Reading a study critically: who, what, compared to what, how selected
  5. The exercise set

Recap: The Toolkit So Far

Stage Tool
Knowledge Types of knowledge; what makes knowledge scientific
Arguments Premises, conclusions, argument structure
Logic Validity, soundness, deduction vs. induction
Failure modes Fallacies, misinformation, persuasion

Today we move from arguments to evidence: what happens when the premises are empirical claims about the world?

One of those failure modes gets the long version today: false cause“after” is not “because.” You learned to name it. Today: how to test it.

Part 1: Hume’s Problem

Why causation is always an inference

What Do We Actually Observe?

2:00 pm — headache. She takes a pill.

3:00 pm — she feels better.

Why did she improve?

Observed and Inferred

  • We observe that Sofia had a headache.
  • We observe that she took a pill.
  • We observe that later she felt better.
  • We do not observe the pill causing the improvement.
  • “The pill worked” is an inference — something we added.

Hume’s point: we observe events and the order they came in. The causal link is something we infer — the word because is not in the observation.

Three Rival Stories

One observation. More than one story that produces it.

The observation is identical under all three. On its own it discriminates between none of them.

And this is n = 1. One episode cannot separate these stories. Science is the problem of finding comparisons that can.

The Missing Comparison

We cannot run both worlds for the same Sofia. The causal effect is a difference we can never directly observe.

Causal effect = what happened minus what would have happened otherwise. Everything that follows is about finding a credible stand-in for that second line.

Hume’s Challenge

“All reasonings concerning matter of fact seem to be founded on the relation of Cause and Effect… causes and effects are discoverable, not by reason but by experience.”

— Hume, Enquiry, Section IV

Three claims nobody has ever seen happen:

  • “Social media causes polarization” — or do polarized people seek out the feeds that already agree with them?
  • “Private schools make better students” — or do they admit the students who were already ahead?
  • “Smoking causes cancer” — true. And still nobody ever watched it happen: that arrow was earned by comparison, not by sight.

Each time we saw a pattern — a constant conjunction, in Hume’s phrase. The word causes is an inference — something we added.

Consequence: the pattern is real; the story we tell about it might not be.

Part 2: Correlation vs. Causation

The four default suspects behind every pattern

What Is a Correlation?

Two variables are correlated when they move together:

  • Ice cream sales rise; drowning deaths rise.
  • Countries that consume more chocolate win more Nobel Prizes.
  • Students who take notes by hand get better grades.

A correlation is a fact about the data.

Any Correlation Has Four Default Suspects

X is correlated with Y. Why? Round up the usual four before looking further:

  1. X causes Y — the story we usually jump to
  2. Y causes X — reverse causation
  3. Z causes both — a confounder (lurking variable)
  4. Chance — coincidence, especially when you look at enough things

Suspect 2: Reverse Causation

  • “Happier employees are more productive.” Or do productive employees get promoted, paid, praised — and become happier?
  • “Police presence is correlated with crime.” Do police cause crime — or are police sent where crime is?

The test: point the arrow both ways, and say each story out loud.

  • X → Y — does that story make sense?
  • Y → X — does that one make sense too?

If both stories work, the correlation cannot tell you which one is true.

Suspect 3: Confounders

Ice cream sales and drownings rise together. Does ice cream cause drowning?

Summer causes both: heat drives ice cream sales and swimming.

The same shape, in claims you might actually believe:

  • “Countries with more McDonald’s have longer life expectancy.”wealth produces both.
  • “People who watch TV news vote more often.”age produces both.
  • “People who take vitamins live longer.” — the kind of person who takes vitamins also exercises and sees a doctor.

Test: ask “what third factor could produce both?”

Suspect 4: Chance

If everyone here flips a coin five times, someone will get five heads. That person has no gift — there were just enough of you.

Data works the same way:

  • US cheese consumption tracks deaths by bedsheet entanglement.
  • Nicolas Cage films per year track swimming-pool drownings.

Nobody noticed these. A computer trawled thousands of datasets and printed the pairs that happened to line up.

Chocolate and Weddings

One country, fourteen years. The two lines move together almost perfectly.

Nobody thinks chocolate causes weddings. So where did this chart come from?

Twenty Countries

That chart was Country 13. The same two variables, in nineteen other countries, show nothing at all.

Look in enough places and one of them will line up. Test: ask “how many were compared before this one was shown to me?”

Part 3: The Inference Audit

A tool for testing every “therefore”

The Problem With “Therefore”

Arguments about evidence hide their weakest step behind inferential words:

therefore — thus — this shows — this proves — this rules out — this is inconsistent with — this confirms

These words assert that the evidence supports the conclusion. They do not demonstrate it.

You already stop at “therefore”: conclusion? premises? unstated assumption? With evidence, that third question has a standing answer — “and no rival explanation fits.” The audit is that check, grown.

The Inference Audit: every time you meet one of these words, stop and test the logic. Six questions, one rule.

Fits Both Stories

Claim: the pill cured Sofia’s headache.

Evidence: she took the pill, and an hour later the headache was gone.

  • Story 1: the pill worked.
  • Story 2: the headache would have gone anyway.

The evidence fits both. On its own it cannot pick one.

Good evidence does not merely fit an explanation — it helps us choose between competing explanations. That is the whole job of the six questions that follow.

The Six Audit Questions

  1. What is the conclusion? What is the author trying to prove, reject, or explain?
  2. What are the premises? What facts or assumptions serve as evidence?
  3. What would each hypothesis predict? Write down what the main hypothesis and the rival would each lead us to expect — before judging.
  4. Does the evidence match the hypothesis being rejected? If yes, the argument has refuted nothing.
  5. Does the evidence discriminate? Or would the same pattern appear under several explanations?
  6. Could the pattern have happened anyway? Ceiling effects, regression to the mean, selection, composition change, shared time trends.

Worked Example: The Cafeteria

Twelve students got sick after Tuesday’s lunch. A new chicken dish had gone on the menu that week.

The cafeteria manager: “Impossible. This kitchen has served food for twenty years without a single case of food poisoning. The sickness only started this month. Therefore, the chicken is not to blame.”

Sounds like a solid defence. Run the audit.

Worked Example: Running the Audit

  1. Conclusion: the new chicken dish is not to blame.
  2. Premises: twenty years with no cases; the sickness began this month.
  3. What would the rival predict? If the new dish caused the outbreak, we would expect… no cases before it appeared, and cases starting once it did.
  4. Does the evidence match the hypothesis being rejected? Yes — exactly.

The manager’s twenty clean years are precisely what “the new dish did it” predicts. The “therefore” is empty.

Worked Example: The Lesson

%%{init: {"theme":"base","flowchart":{"useMaxWidth":true,"curve":"basis","nodeSpacing":45,"rankSpacing":75},"themeVariables":{"fontSize":"20px"}}}%%
flowchart LR
  A["<b>Story 1</b><br/>The new chicken<br/>dish did it"] --> E
  B["<b>Story 2</b><br/>Something else new<br/>did it"] --> E
  E["<b>What the manager saw</b><br/>20 clean years,<br/>then 12 students sick"]
  style A fill:#f9fafb,stroke:#64748b,stroke-width:1.5px,color:#1e293b
  style B fill:#f9fafb,stroke:#64748b,stroke-width:1.5px,color:#1e293b
  style E fill:#1e293b,stroke:#1e293b,color:#f9fafb
  linkStyle 0,1 stroke:#b44527,stroke-width:2.5px

Both stories predict exactly what he saw. His evidence cannot choose between them.

The comparison that could: students who ate the chicken vs. those who ate there and did not.

He never asked what the other story would predict — confirmation bias, dressed as data analysis.

Trap 1: Ceiling Effects

The closer you already are to the top, the less room is left to improve.

A student scores 50, then 80, then 90. The first jump is +30. The second is +10.

The mistake: “the new teaching method has stopped working.”

Slower improvement is not a weaker cause. Sometimes there is simply less left to gain.

Trap 2: Regression to the Mean

Extreme results are part real and part luck — and luck does not repeat.

Sofia reached for the pill when the headache was at its worst. Headaches at their worst tend to ease, whatever you do.

The mistake: “the pill worked.”

Measure anything at its extreme, and the next measurement will usually be less extreme. That looks exactly like a treatment working.

Trap 3: Selection

The people who end up in a group are often different before anything happens to them.

Students who sit in the front row get better grades.

The mistake: “sitting at the front raises your grade.”

Ask who ended up in the group, and whether they were already different. Nobody assigned those seats.

Trap 4: Composition Change

An average can move because the group changed, even if nobody in it changed at all.

The mistake: “standards are slipping.”

Not one student got worse. The average fell because ten new students joined — before reading a moving average, ask whether it is still the same group.

The Quick Logic Check

Three questions to remember

For every “therefore” in an argument, ask:

  1. What is the evidence supposed to prove?
  2. What else could produce the same evidence?
  3. What comparison would tell those explanations apart?

Short version: what’s the claim? what else? what would decide?

If both stories predict what we observed, the evidence cannot tell us which story is right.

Part 4: Reading a Study Critically

Who? What? Compared to what? How was it selected?

A Study Is an Argument

A scientific study is not a fact — it is an argument: premises (data, methods) supporting a conclusion (findings).

So everything the course has taught about arguments — plus the Inference Audit — applies. Plus four questions for studies:

  1. Who was studied? — can we generalise from this sample?
  2. What was actually measured? — does the measure capture the thing?
  3. Compared to what? — what tells us the treatment made the difference?
  4. How was the result selected? — the planned test, or the one that looked best after many tries?

Who? What? Compared to what? How was it selected?

Question 1: Who Was Studied?

  • Size: small samples are noisy — chance swings them further, in both directions.
  • Selection: who ended up in the sample, and who never had the chance? An online poll about internet regulation samples… people on the internet.
  • Generalisation: much of psychology studies Western, Educated, Industrialised, Rich, Democratic undergraduates — and concludes about humanity.

Ask: who is in the sample, who is missing, and does the conclusion quietly extend beyond the people actually studied?

Question 2: What Was Actually Measured?

A measure of something is not the thing itself.

  • “Happiness” as an answer on a 1–10 survey scale.
  • “Crime” as the number of arrests — which also measures how hard police were looking.
  • “Learning” as a score on one exam.

The test: if the measure moved, would the thing we actually care about necessarily have moved?

Question 3: Compared to What?

“Crime fell 20% after the policy.”

  • Compared with the year before?
  • Compared with the trend that was already underway?
  • Compared with similar cities that did not adopt the policy?

Without a credible comparison, “after” is not “because of.”

Experiment or Observation?

The comparison has to come from somewhere.

A coin has no opinion about who deserves the treatment, so the two groups end up alike — including on the things you never thought to measure.

Randomisation is not a statistical nicety. It is the closest we get to the row we can never observe.

Question 4: How Was the Result Selected?

Twenty research teams test the same idea. Eighteen find nothing.

The two that found something get published. The other eighteen go in a drawer.

You read the literature, see two positive results, and conclude the effect is real. Nobody lied. Nobody even did anything unusual.

The question to ask: how many analyses were tried before this one appeared? Bias does not require fraud — the filter between what gets run and what gets read is simply not neutral.

And funding? A conflict of interest tells you to look harder at the design and the reporting — never on its own that the finding is false.

Try Enough Things

Suppose there is no real effect at all.

A researcher tests twenty outcomes. Then a few subgroups. Then a few statistical models.

Sooner or later one of them looks impressive. Every individual step felt reasonable.

The result is real in this dataset — which is not the same as evidence about the world.

This has a name: p-hacking — trying many analyses and reporting the ones that came out well. It is Part 2’s chance explanation, inside a single study.

The Exercise Set

Name the Rival. Name the Comparison.

For each claim, the same three questions you will use on everything else: what is it supposed to prove? what else could produce the same pattern? what comparison would tell them apart?

  1. “Students who attend office hours get higher grades. Therefore, attending office hours improves grades.”
  2. “The lowest-ranked schools in 2020 improved the most by 2024. Therefore, the turnaround program works.”
  3. “Students who use AI tools get lower grades. Therefore, AI damages learning.”
  4. “Crime fell 30% after the new mayor took office. This proves her policies work.”

The Exercise Set — Answers

1. “Students who attend office hours get higher grades — therefore office hours improve grades.”

  • Prime suspect: confounder. The conscientiousness that gets you to office hours also gets you to the reading.
  • The comparison: randomly invite half the class. Compare invited to not invited — never attenders to non-attenders.

2. “The lowest-ranked schools in 2020 improved the most by 2024 — therefore the turnaround program works.”

  • Prime suspect: regression to the mean. Ranking last takes real weakness plus bad luck, and luck does not repeat.
  • The comparison: equally low-ranked schools that got no program. Did they rebound just as much?

3. “Students who use AI tools get lower grades — therefore AI damages learning.”

  • More than one rival here. Selection: students already struggling reach for the tool. Reverse causation: falling behind causes the AI use, not the other way round. Measurement: a grade measures what the assessment rewards.
  • The comparison: similar students randomly given access or not — or the same students before and after a rule that changed access.

4. “Crime fell 30% after the new mayor took office — this proves her policies work.”

  • Prime suspect: a shared time trend, plus regression to the mean: she was elected right after a spike.
  • The comparison: the same four years in comparable cities that did not change mayor.

Every answer has the same shape: name the rival, then name the comparison that would tell them apart.

Some claims have several rivals at once. That is normal — the question is never which fallacy is this, it is whether the evidence can choose.

Wrap-Up

  • Causation is never observed — only inferred (Hume). A causal effect is a difference from a world we do not get to see.
  • Every correlation has four default suspects — and some patterns need no cause at all. The burden is to discriminate, not to fit.
  • The Inference Audit: derive what the rival predicts before accepting any “therefore.”
  • A study is an argument: who was studied, what was measured, compared to what, how the result was selected.

And the version to carry out of here: what’s the claim? what else? what would decide?

References

  • Hume, D. (2007). An enquiry concerning human understanding (P. Millican, Ed.). Oxford University Press. (Original work published 1748; Section IV.)