Logic and Scientific Thinking

Correlation, Causation

Bogdan G. Popescu

Tecnologico de Monterrey

Welcome

Learning Outcomes

Overview

  1. Distinguish correlation from causation.
  1. Evaluate the credibility of sources and the strength of evidence behind claims.
  1. Read a scientific study critically: sample size, methodology, and potential biases.

Roadmap

Where we are going today

  1. Hume’s problem: why causal knowledge is never directly observed
  2. Correlation vs. causation: the four explanations for any pattern
  3. The Inference Audit: a tool for testing every “therefore”
  4. Reading a study critically: samples, methods, biases, p-hacking
  5. The exercise set

Recap: The Toolkit So Far

Stage Tool
Knowledge Types of knowledge; what makes knowledge scientific
Arguments Premises, conclusions, argument structure
Logic Validity, soundness, deduction vs. induction
Failure modes Fallacies, misinformation, persuasion

Today we move from arguments to evidence: what happens when the premises are empirical claims about the world?

Part 1: Hume’s Problem

Why causation is always an inference

What Do We Actually Observe?

Drop a piece of chalk. What do you see?

  • You see the hand release the chalk.
  • You see the chalk fall.
  • You do not see the releasing cause the falling.

Hume (1748): we never observe causation itself — only constant conjunction: A happens, then B happens, again and again.

Hume’s Challenge

“All reasonings concerning matter of fact seem to be founded on the relation of Cause and Effect… causes and effects are discoverable, not by reason but by experience.”

— Hume, Enquiry, Section IV

Every causal claim — “smoking causes cancer,” “austerity caused the recession,” “social media causes polarization” — is an inference from observed patterns, not an observation.

Consequence: causal claims can always be wrong. The pattern is real; the story we tell about it might not be.

Part 2: Correlation vs. Causation

Four stories behind every pattern

What Is a Correlation?

Two variables are correlated when they move together:

  • Ice cream sales rise; drowning deaths rise.
  • Countries that consume more chocolate win more Nobel Prizes.
  • Students who take notes by hand get better grades.

A correlation is a fact about the data.

Any Correlation Has (At Least) Four Explanations

X is correlated with Y. Why?

  1. X causes Y — the story we usually jump to
  2. Y causes X — reverse causation
  3. Z causes both — a confounder (lurking variable)
  4. Chance — coincidence, especially with small samples

Explanation 2: Reverse Causation

  • “Happier employees are more productive.” Or do productive employees get promoted, paid, praised — and become happier?
  • “Police presence is correlated with crime.” Do police cause crime — or are police sent where crime is?

Test: ask which way the arrow could plausibly run. If both directions tell a coherent story, the correlation alone cannot decide.

Explanation 3: Confounders

Ice cream sales and drownings rise together. Does ice cream cause drowning?

Summer causes both: heat drives ice cream sales and swimming.

Classic confounders in social science:

  • Wealth: correlates with health, education, trust, institutions…
  • Age: correlates with income, political views, media habits…
  • Education: correlates with almost everything

Test: ask “what third factor could produce both?”

Explanation 4: Chance

With enough variables, absurd correlations are guaranteed:

  • US cheese consumption correlates with deaths by bedsheet entanglement.
  • Nicolas Cage films per year correlate with swimming pool drownings.

If you test 100 variable pairs, roughly 5 will look “statistically significant” by pure luck.

This is not a joke about bad researchers — it is a mathematical property of searching for patterns. Remember it when we reach p-hacking.

The Ecological Fallacy

A pattern in groups need not hold for the individuals in them:

  • Across 1930 US states: more foreign-born residents, higher English-literacy rates.
  • Yet foreign-born individuals were less literate — they settled in high-literacy states (Robinson 1950).
  • Richer states lean to one party; richer voters lean to the other.
  • Test: match the data’s level (state, school, country) to the claim’s level.

Part 3: The Inference Audit

A tool for testing every “therefore”

The Problem With “Therefore”

Arguments about evidence hide their weakest step behind inferential words:

therefore — thus — this shows — this proves — this rules out — this is inconsistent with — this confirms

These words assert that the evidence supports the conclusion. They do not demonstrate it.

The Inference Audit: every time you meet one of these words, stop and test the logic. Six questions, one rule.

The Core Mistake It Catches

Most bad empirical arguments share one structure:

  1. The author has a favorite hypothesis, H.
  2. The evidence is consistent with H.
  3. The author concludes H is true — without checking whether the evidence is also consistent with the rival hypothesis.

The rule: Do not ask only whether the evidence fits the author’s story. Ask whether it also fits the story the author is trying to reject.

The Six Audit Questions

  1. What is the conclusion? What is the author trying to prove, reject, or explain?
  2. What are the premises? What facts or assumptions serve as evidence?
  3. What would each hypothesis predict? Write down what the main hypothesis and the rival would each lead us to expect — before judging.
  4. Does the evidence match the hypothesis being rejected? If yes, the argument has refuted nothing.
  5. Does the evidence discriminate? Or would the same pattern appear under several explanations?
  6. Could the pattern arise mechanically? Ceiling effects, mean reversion, selection, composition change, general time trends.

Worked Example: Regional Convergence

A passage from a realistic policy report:

“Some blame federalism for the slow economic convergence of poor regions. But this cannot be right: convergence was fastest under the earlier centralized system and slowed down after decentralization. Therefore, federalism does not explain slow convergence.”

Sounds rigorous. Run the audit.

Worked Example: Running the Audit

  1. Conclusion: federalism does not explain slow convergence.
  2. Premises: convergence was fast under centralization, slow after decentralization.
  3. What would the rival predict? If federalism constrained convergence, we would expect… fast convergence under centralization and slower convergence after decentralization.
  4. Does the evidence match the hypothesis being rejected? Yes — exactly.

The evidence the author uses to refute the federalism hypothesis is precisely what the federalism hypothesis predicts. The “therefore” is empty.

Worked Example: The Lesson

  • The author never asked what the rival hypothesis would predict.
  • The evidence was consistent with both stories — so it discriminates between neither.
  • The author needed a better comparison (regions with more vs. less autonomy?) or a weaker claim.

Name the mistake: failure to derive the rival hypothesis’s prediction.

This is confirmation bias, presented as data analysis.

Question 6: Mechanical Explanations

Before accepting a substantive story, check whether the pattern could arise from arithmetic alone:

Pattern Possible mechanical cause
“Improvement slowed after the reform” Ceiling effect, diminishing returns
“The worst performers improved the most” Mean reversion
“Program participants did better” Selection: who joins?
“Average scores fell as enrollment grew” Composition change
“X rose over the decade, and so did Y” Both follow a general time trend

None of these requires the author’s story to be true — or false. They require the author to rule them out.

The Quick Logic Check

The version to memorize

For every “therefore” in an argument, ask:

  1. What is the claim?
  2. What is the evidence?
  3. What would each competing explanation predict?
  4. Does the evidence fit the author’s explanation, the rival, both, or neither?
  5. Does the conclusion actually follow?

Final test: if the rival hypothesis predicts the same pattern we observe, the evidence cannot be used to reject that rival.

Part 4: Reading a Study Critically

Sample, method, bias

A Study Is an Argument

A scientific study is not a fact — it is an argument: premises (data, methods) supporting a conclusion (findings).

So everything the course has taught about arguments — plus the Inference Audit — applies. Plus three study-specific questions:

  1. Who was studied? (sample)
  2. How was the effect measured? (methodology)
  3. Who benefits from the result? (bias and incentives)

Question 1: The Sample

  • Size: a survey of 40 students tells you little; “n = 12” is a warning sign.
  • Selection: an online poll about internet regulation samples… people on the internet.
  • WEIRD samples: much of psychology studies Western, Educated, Industrialized, Rich, Democratic populations — mostly undergraduates — and generalizes to humanity.

Ask: who is in the sample, who is missing, and does the conclusion quietly extend beyond the people actually studied?

Question 2: The Methodology

  • Experiment or observation? Random assignment breaks confounding; observation does not.
  • What is actually measured? “Happiness” measured as a 1–10 survey answer is not the same as happiness.
  • Compared to what? “Crime fell after the policy” — compared to cities without the policy? To the previous trend?

Question 3: Bias and Incentives

  • Who funded the study?
  • Would the researcher’s career benefit from one result more than another?
  • Journals prefer positive, surprising findings — boring true results are harder to publish than exciting fragile ones.

Bias does not mean fraud. It means the filter between all studies conducted and the studies you get to read is not neutral.

p-Hacking

Recall: test 100 comparisons, and ~5 look significant by chance.

Now imagine a researcher who tries many outcomes, subgroups, and model specifications — and reports the one that “worked.”

  • Each individual step can feel reasonable.
  • The reported result is real in the data — and meaningless about the world.
  • The published paper shows you the 1 comparison, not the 99.

This is the chance explanation from Part 2, at scale.

The Exercise Set

The Exercise Set: Audit Four Claims

For each claim: identify the pattern, list the rival explanations (reverse causation, confounder, chance, mechanical).

  1. “Students who attend office hours get higher grades. Therefore, attending office hours improves grades.”
  2. “The lowest-ranked schools in 2020 improved the most by 2024. Therefore, the turnaround program works.”
  3. “Countries that adopted austerity grew slower. Therefore, austerity causes slow growth.”
  4. “Crime fell 30% after the new mayor took office. This proves her policies work.”

Wrap-Up

  • Causation is never observed — only inferred (Hume).
  • Every correlation has four candidate explanations; the burden is to discriminate, not to fit.
  • The Inference Audit: derive what the rival predicts before accepting any “therefore.”
  • A study is an argument: check sample, method, incentives.
  • Single studies are weak evidence; replication and convergence are strong.

References

  • Hume, D. (2007). An enquiry concerning human understanding (P. Millican, Ed.). Oxford University Press. (Original work published 1748; Section IV.)
  • Robinson, W. S. (1950). Ecological correlations and the behavior of individuals. American Sociological Review, 15(3), 351–357.