Logic and Scientific Thinking

Reading a Study Critically — Six Moves for Any Social Science Article

Bogdan G. Popescu

Tecnológico de Monterrey

Before We Start: The Final

  • Thursday, September 10, in class — closed book, individual.
  • Covers Sessions 10–14, today included. Not cumulative.
  • The required readings for those sessions are examinable.

Today is the last session — and the six moves are the whole toolkit, used at once.

The Core Idea

An article is not a collection of facts.

It is an argument: evidence and assumptions, assembled to make you accept a conclusion.

Facts you memorise. Arguments you evaluate.

What You Will Be Able to Do

By the end of today, with any article in any social science:

  • State precisely what you are being asked to believe
  • Find the assumption the authors never wrote down
  • Say what else could have produced their result
  • Give a calibrated verdict instead of a verdict of taste

The Six Moves

%%{init: {"flowchart": {"useMaxWidth": true, "htmlLabels": false, "nodeSpacing": 14, "rankSpacing": 34, "padding": 16}, "themeVariables": {"fontSize": "14px", "edgeLabelBackground": "#ffffff"}}}%%
flowchart LR
  M1["1<br/>CLAIM"] --> M2["2<br/>ARGUMENT"] --> M3["3<br/>MEASURE"] --> M4["4<br/>COMPARE"] --> M5["5<br/>RIVALS"] --> M6["6<br/>VERDICT"]
  style M1 fill:#1e293b,color:#f9fafb,stroke:#1e293b
  style M2 fill:#1e293b,color:#f9fafb,stroke:#1e293b
  style M3 fill:#4a7c6f,color:#f9fafb,stroke:#4a7c6f
  style M4 fill:#4a7c6f,color:#f9fafb,stroke:#4a7c6f
  style M5 fill:#b44527,color:#f9fafb,stroke:#b44527
  style M6 fill:#b7943a,color:#f9fafb,stroke:#b7943a

Same six, every time, in this order. They are a procedure, not a vocabulary list.

Move 1 — The Claim

Four Kinds of Empirical Claim

  • Descriptive — how the world is
  • Causal effect — X changed Y
  • Mechanismwhy X changed Y
  • Generalization — it holds elsewhere too

Different claims need different evidence. Papers slide between them.

One Topic, Four Claims

Descriptive Democracies rarely fight each other.
Causal Democracy makes states less likely to fight.
Mechanism …because leaders fear electoral punishment.
General Any democracy, anywhere, will be peaceful.

Evidence for the first supports none of the other three.

Move 1 in Practice

Ask, before anything else:

  • Which of the four is this paper’s main claim?
  • Which does the title promise?
  • Which does the conclusion section deliver?

When those three disagree, you have already found something.

Move 2 — The Argument

Standard Form

Strip the prose. Number the premises. Put the conclusion last.

“Countries that adopted the reform grew faster afterwards. The reform works.”

P1. Adopters grew faster after adopting.
C. The reform causes growth.

One premise, and a very large conclusion. Something is missing.

The Missing Premise

P1. Adopters grew faster after adopting.
P2. (unstated) Adopters and non-adopters were otherwise comparable.
C. The reform causes growth.

P2 appears nowhere in the text — and the argument collapses without it.

An unstated premise doing real work is an enthymeme. Nearly every paper has one.

Read It Charitably

Before attacking, build the strongest version the text supports.

  • Fill in the premise the author would actually endorse
  • Not the silliest one that makes them wrong
  • If your version is easy to demolish, you built it wrong

A critique of a weak reconstruction tells the reader nothing.

Where Hidden Premises Hide

Watch the words that open or close a category:

  • such as, for example, including — opens
  • these, any of these, the — closes
  • mere, sufficient, only, cannot, rules out — claims a limit

A paper that opens a category and then closes it a sentence later has assumed the list was complete.

That assumption is almost never stated.

Move 3 — Measurement

What Was Actually Observed?

The concept in the title is rarely the thing in the dataset.

Concept claimed Thing measured
democracy an index score, 0–10
state capacity tax revenue over GDP
wellbeing a 1–7 survey answer
prejudice …often, a vote

Ask: does the measure capture the concept, or merely correlate with it?

Two Levels

%%{init: {"flowchart": {"useMaxWidth": true, "htmlLabels": false, "nodeSpacing": 26, "rankSpacing": 60, "padding": 20}, "themeVariables": {"fontSize": "15px", "edgeLabelBackground": "#ffffff"}}}%%
flowchart LR
  G["MEASURED<br/>groups: countries,<br/>regions, towns"] -.->|"does this follow?"| I["CLAIMED<br/>individuals:<br/>voters, people"]
  style G fill:#4a7c6f,color:#f9fafb,stroke:#4a7c6f
  style I fill:#b44527,color:#f9fafb,stroke:#b44527

Regions with more immigrants voted further right.

It does not follow that immigrants, or anyone near them, voted further right.

Group data, individual claim: the ecological fallacy.

The Diagnostic

Two questions, always paired:

At what level was the data measured?

At what level is the claim being made?

Group evidence licenses group conclusions. Nothing more.

Move 4 — Comparison

Compared to What?

“Crime fell 12% after the policy.”

  • Compared with the year before? (what else changed?)
  • With cities that had no policy? (were they similar?)
  • With the existing downward trend? (was it already falling?)

A number without a comparison is not evidence for anything.

Find the Assumption Sentence

Every design buys its comparison with an assumption. Good papers write it down.

Design The assumption
Before/after nothing else changed at the same time
Treated vs. control both groups were on parallel paths
Instrument it affects the outcome only through the treatment
Controls/matching no unmeasured confounder is left
Experiment assignment really was random

Then Ask One Question

Once you have found the sentence:

Do I believe it?

Not “is it sophisticated.” Not “did it get published.”

Can you name something concrete that would make it false?

Move 5 — Rivals

Generate Rivals First

Before judging the evidence, write down what each explanation predicts.

“Students who attended tutoring earned higher grades. Tutoring works.”

  • Selection — motivated students choose tutoring
  • Reverse causation — struggling students are sent to tutoring
  • Mean reversion — extreme scores drift back anyway
  • Something else changed — new syllabus, easier exam

The Discriminating Test

Ask of the evidence:

Does it discriminate? Or would the same pattern appear under several explanations?

The rule: if a rival predicts the same pattern, that evidence cannot reject the rival.

The burden is to discriminate, not merely to fit.

Effect Is Not Explanation

X caused Y and we know why X caused Y are two separate arguments.

A design can nail the first while leaving several mechanisms alive.

Papers win the first and then write the conclusion as if they had won both.

The Mechanical Checklist

Before crediting any explanation, rule out the boring ones:

  • Selection — who entered the sample, and how
  • Mean reversion — extremes drift back
  • Composition change — the group itself changed
  • Time trends — it was already moving
  • Ceiling/floor — no room left to move

Move 6 — Verdict

Three Tiers, Not a Score

Well supported — the evidence carries this

Plausible but not established — consistent, not discriminated

Not established by these data — claimed beyond the evidence

Most papers land in all three at once. Say which is which.

A Flaw Is Not a Refutation

Finding a weakness does not show the conclusion is false.

  • A bad argument can have a true conclusion
  • “Their design has a limitation” ≠ “they are wrong”
  • The verdict is about what the evidence carries

Your job is to locate the line between what was shown and what was said.

The Exercise

The Paper

Dinas, Matakos, Xefteris & Hangartner (2019), Political Analysis.

  • Setting: Greek islands, 2015 refugee crisis
  • Design: islands near Turkey received refugees; distant islands did not
  • Data: vote shares, 95 municipalities, before and after
  • Finding: +2 points for the far-right party Golden Dawn

The design is strong. Two strategies, placebo tests, a dose-response gradient.

What They Conclude

Three sentences from the paper:

“mere exposure is sufficient to fuel prejudice and change political behavior” (p. 247)

competition over “scarce resources such as access to jobs, housing, or education… there is no specific competition… over any of these resources” (p. 253)

“the ensuing chaos on affected islands” (p. 253)

Your Task

15 minutes, in pairs. Run the moves you can:

  • Move 1 — which kind of claim is “fuels prejudice”?
  • Move 2 — reconstruct the resource argument; what is unstated?
  • Move 3 — what was measured? at what level?
  • Move 5 — what else would produce a 2-point rise?

Then give a three-tier verdict.

Reveal 1 — The Hidden Premise

P1. Group conflict needs competition over scarce resources.
P2. No competition over jobs, housing, or education.
P3. (unstated) Those three are the relevant scarce resources.
C. So group conflict does not explain the backlash.

such as” opens the category. “any of these” closes it. P3 is never defended.

Policing, transport, clinics, public space — also scarce. The authors’ own phrase: “ensuing chaos.”

Reveal 2 — Two Jumps

%%{init: {"flowchart": {"useMaxWidth": true, "htmlLabels": false, "nodeSpacing": 26, "rankSpacing": 50, "padding": 18}, "themeVariables": {"fontSize": "14px", "edgeLabelBackground": "#ffffff"}}}%%
flowchart LR
  V["MEASURED<br/>town vote shares"] -.->|"group to individual"| I["individual<br/>voters"]
  I -.->|"votes to minds"| P["CLAIMED<br/>prejudice"]
  style V fill:#4a7c6f,color:#f9fafb,stroke:#4a7c6f
  style I fill:#f9fafb,color:#334155,stroke:#64748b
  style P fill:#b44527,color:#f9fafb,stroke:#b44527

Prejudice is never measured. The outcome is a vote share.

They measured votes, not prejudice.

Reveal 3 — Rivals

%%{init: {"flowchart": {"useMaxWidth": true, "htmlLabels": false, "nodeSpacing": 20, "rankSpacing": 66, "padding": 18}, "themeVariables": {"fontSize": "14px", "edgeLabelBackground": "#ffffff"}}}%%
flowchart LR
  A["A. Mere exposure<br/>seeing refugees"] --> R["Golden Dawn<br/>vote rises"]
  B["B. Disruption<br/>chaos, strained services"] --> R
  C["C. Salience<br/>immigration dominates news"] --> R
  style A fill:#b44527,color:#f9fafb,stroke:#b44527
  style B fill:#f9fafb,color:#334155,stroke:#64748b
  style C fill:#f9fafb,color:#334155,stroke:#64748b
  style R fill:#1e293b,color:#f9fafb,stroke:#1e293b

All three predict the same result. The evidence does not discriminate.

The Calibrated Verdict

Well supported: refugee arrivals raised Golden Dawn’s vote share by ~2 points.

Plausible, not established: that seeing refugees did it, rather than the disruption around them.

Not established by these data: that anyone’s prejudice changed.

This is a good paper. Notice that the verdict still has three tiers.

The Handout

Six Moves

  1. CLAIM — what exactly am I asked to believe?
  2. ARGUMENT — P1, P2, hidden P3, C?
  3. MEASUREMENT — what was measured, at what level?
  4. COMPARISON — compared with what, on what assumption?
  5. RIVALS — what else predicts this? does it discriminate?
  6. VERDICT — supported / plausible / not established

One Last Question

After every verdict, ask:

What additional evidence would change my mind?

If nothing would, you are not analysing. You are just disagreeing.