Latest Trending Discover Timelines Categories
←All explainers

Technology explainer

Why Doesn't a Correlation Between Two Things Mean One Causes the Other?

Correlation describes variables moving together; causation predicts what would change under an intervention. Reverse causation, confounding, selection, measurement, and chance can create associations, so causal confidence depends on study design, credible comparisons, time order, replication, and converging evidence.

A correlation means two measured variables change together more often than expected by chance. It does not by itself show that changing one variable will change the other. The relationship may reflect direct causation, reverse causation, a shared cause, biased sampling, measurement choices, coincidence, or a mixture of several mechanisms.

The distinction in 30 seconds

  • Correlation describes a pattern. Higher values of one variable tend to accompany higher or lower values of another.
  • Causation describes an intervention. If one factor were deliberately changed while relevant alternatives were held comparable, the outcome would change because of it.
  • Timing helps but is insufficient. A cause must precede its effect, yet many earlier variables merely predict later outcomes.
  • Statistical adjustment relies on assumptions. Researchers can control measured confounders, not every unknown or poorly measured factor.
  • Evidence comes in layers. Randomisation, natural experiments, repeated findings, mechanism, and dose-response patterns can strengthen a causal case.

What exactly is correlation?

Correlation summarises how two variables vary together. A positive correlation means larger values of one tend to accompany larger values of the other. A negative correlation means larger values tend to accompany smaller values. A correlation near zero means no clear linear relationship, although a curved or threshold relationship may still exist.

The coefficient does not explain why the pattern exists. It also does not tell the reader whether the effect is large enough to matter, whether the measurements are reliable, or whether the relationship applies outside the studied group.

What would a causal claim mean?

A causal question compares potential outcomes. What would happen to the same person, school, company, patient, or system if the exposure changed while everything else relevant remained comparable? In reality, we cannot observe the same unit simultaneously in both conditions. Study design tries to construct a credible comparison group that represents the unobserved alternative.

This is why causal inference is difficult. A statement such as “people who do X have better outcomes” compares different people. A causal statement such as “making a person do X improves the outcome” asks what would happen after an intervention.

Six reasons a correlation can appear

Explanation Pattern Example structure
Direct causation X changes Y A treatment lowers a biological measure
Reverse causation Y changes X Early illness reduces physical activity rather than inactivity causing every observed symptom
Confounding A third factor changes both X and Y Income affects access to a technology and educational outcomes
Selection bias Who enters or remains in the study creates the link Only highly motivated users complete both a programme and its follow-up test
Measurement or analysis artefact Definitions, errors, or model choices generate or distort the association Self-reported exposure shares bias with a self-reported outcome
Chance and multiple testing A random pattern appears among many comparisons One “significant” result emerges after testing dozens of unrelated outcomes

Real studies can contain several at once. A modest causal effect may be exaggerated by confounding and weakened by measurement error.

Why is a confounder more than “another variable”?

A confounder is related to both the exposure and outcome and is not simply a consequence of the exposure. It creates an alternative path connecting them. Age, prior health, education, geography, income, and baseline ability are common candidates, but the relevant confounders depend on the question.

Adding variables to a regression model does not guarantee removal of confounding. A factor may be measured crudely, omitted, or modelled with the wrong relationship. Researchers can also adjust for the wrong variable, including one caused by the exposure, and thereby hide part of the real effect or create bias.

How does reverse causation fool a longitudinal study?

Following people over time establishes order better than measuring everything once. But an undetected early stage of the outcome can influence the exposure before the formal outcome is recorded. People who feel cognitive decline may alter device use, patients with developing disease may reduce exercise, and struggling companies may adopt a technology in response to poor performance.

Researchers can use baseline screening, lagged analyses, repeated measurements, or exclusion of early outcomes, but those steps reduce rather than eliminate the possibility. Longitudinal means “over time,” not automatically “causal.”

What do statistical controls actually do?

Matching, stratification, regression, weighting, and related methods compare units that look similar on measured characteristics. Their credibility depends on whether the important confounders were identified, measured accurately, and modelled appropriately. Large datasets improve precision but do not repair a systematically biased comparison.

A small p-value answers a limited question about compatibility with a statistical model under assumptions. It does not measure the probability that the hypothesis is true, prove causation, or show practical importance. Confidence intervals and effect sizes reveal more about magnitude and uncertainty.

Why randomisation is powerful

In a well-designed randomized controlled trial, chance assigns participants to intervention groups. Before the intervention, known and unknown background factors should be balanced on average, so a later outcome difference can more credibly be attributed to the assigned intervention.

Randomisation is not magic. Non-adherence, dropout, unblinded behaviour, small samples, outcome switching, contamination between groups, and poor measurement can undermine the trial. Some questions are too dangerous, unethical, expensive, or long-term to randomise.

What can researchers use when a trial is impossible?

  • Natural experiments: a policy, lottery, threshold, or external event creates comparison groups not chosen by the participants.
  • Instrumental variables: a factor shifts exposure but is assumed to affect the outcome only through that exposure.
  • Regression discontinuity: units just above and below a decision cutoff are compared.
  • Difference-in-differences: outcome changes in an affected group are compared with changes in an unaffected group.
  • Sibling, family, or fixed-effects designs: stable shared characteristics are partly controlled by comparing within groups.
  • Negative controls: an exposure or outcome that should not be causally affected helps detect residual bias.

Each method replaces randomisation with assumptions that must be defended. No statistical label automatically makes an observational analysis causal.

How does a causal case become convincing?

Confidence grows when different evidence points in the same direction: the cause precedes the effect, the association is reproducible, larger exposure produces a plausible response, reducing exposure changes the outcome, a mechanism exists, alternative explanations are tested, and distinct research designs reach compatible estimates.

Not every causal relationship shows a simple dose response, and a known mechanism is not required before an effect can be real. These features are considerations, not a checklist that mechanically converts association into proof.

A practical way to read an observational headline

  1. Identify the exact exposure and outcome. Replace broad words such as “screens,” “AI,” or “health” with what was actually measured.
  2. Ask who was studied. Note age, country, selection, exclusions, sample size, and follow-up loss.
  3. Check measurement quality. Device readings, administrative records, clinical tests, and self-reports have different errors.
  4. Look for baseline differences. Were groups already different before the exposure or outcome?
  5. List plausible common causes. Which variables might influence both sides?
  6. Check time order and reverse causation. Could the emerging outcome have changed the exposure?
  7. Read the effect size and uncertainty. Do not stop at “significant.”
  8. Compare the conclusion with the design. Words such as “linked,” “associated,” or “predicted” should not silently become “caused.”

The screen-time example

NewTqnia covered an eight-year Finnish study in which greater cumulative screen time was associated with better attention, learning, and working-memory scores among 260 adolescents: More Childhood Screen Time Was Linked to Sharper Teen Thinking.

The finding is a real association in that sample, but it does not show that assigning more screen time would improve cognition. The study did not separate educational work, games, communication, and passive viewing. Children with stronger baseline skills or different family resources may choose different screen activities. Self-reported exposure and attrition may also shape the result. The correct next question is what kind of use, for whom, under which conditions, has what causal effect.

“Not causal” does not mean “useless”

Observational evidence can detect harms, reveal rare outcomes, generate hypotheses, estimate real-world patterns, and answer questions that cannot be randomised. The disciplined conclusion is neither “correlation proves causation” nor “correlation means nothing.” It is a calibrated statement of what the design supports and what additional evidence would change confidence.

The mental model to remember

Correlation is a map of observed co-movement. Causation is a claim about what would change under intervention. Moving from one to the other requires a credible comparison that blocks alternative paths, establishes time order, measures the relevant variables, and survives tests using different data and designs.

First appeared in

More Childhood Screen Time Was Linked to Sharper Teen Thinking

A new version of NewTqnia is ready.