Module 3 — Correlation vs Causation

Why relationships in healthcare data are rarely straightforward

When two things move together, does one cause the other?

Healthcare analytics frequently identifies relationships between variables.

For example:

  • higher deprivation and worse health outcomes
  • frailty and increased admissions
  • obesity and diabetes
  • delayed discharge and longer Length of Stay
  • smoking and stroke incidence

When two variables appear related, this is known as a correlation.

However, correlation alone does not prove that one variable directly causes the other.

This distinction matters enormously in healthcare analytics because healthcare systems are complex, interconnected and influenced by many overlapping factors.

What Is Correlation?

Imagine we observe:

Area Fast-food density Obesity
Area A Low Low
Area B High High

We might conclude:

fast-food outlets directly cause obesity

But what else may matter?

  • income
  • access to healthy food
  • physical activity
  • deprivation
  • housing
  • education

The relationship may be real — but the pathway behind it is often more complex than it first appears.

Correlation describes a relationship between two variables.

For example:

Observation Relationship
Areas with higher deprivation often have higher stroke mortality Positive correlation
Increased physical activity is associated with lower cardiovascular risk Negative correlation

Correlation helps us:

  • identify patterns
  • detect variation
  • generate hypotheses
  • prioritise investigation

But correlation alone cannot tell us:

  • why the relationship exists
  • whether the relationship is direct
  • whether other factors are involved

Correlation Does Not Automatically Mean Causation

Suppose we observe:

Area Stroke Mortality
Most deprived areas Higher
Least deprived areas Lower

It may be tempting to conclude:

“Deprivation directly causes higher stroke mortality.”

However, healthcare outcomes are rarely driven by a single factor.

Deprivation may instead be linked to:

  • higher smoking prevalence
  • hypertension
  • obesity
  • diabetes
  • delayed healthcare access
  • poorer housing
  • reduced access to prevention
  • increased multimorbidity
  • environmental factors

In reality, deprivation often acts as a marker for a wider set of underlying risks and system pressures.

Example — Stroke Outcomes and Inequalities

Stroke outcomes often vary significantly between populations.

For example, more deprived populations may experience:

  • higher stroke incidence
  • earlier onset of stroke
  • increased mortality
  • longer rehabilitation needs
  • poorer long-term outcomes

However, these differences are usually influenced by multiple interconnected factors including:

  • prevalence of hypertension
  • smoking rates
  • access to primary care
  • anticoagulation gaps in atrial fibrillation
  • housing and social conditions
  • delayed presentation to services
  • rehabilitation access
  • transport and geography

This means:

the observed inequality is real, but the pathway causing it is complex.

The example below illustrates how healthcare inequalities often emerge through multiple interconnected risk factors rather than a single direct cause-and-effect relationship.

Poor Interpretation vs Better Interpretation

Observed:

higher deprivation → higher stroke mortality

Poor interpretation:

“Deprivation causes stroke.”

Better interpretation:

“Deprivation is associated with a set of behavioural, environmental, biological and access-related risks that may contribute to stroke outcomes.”

Confounding Variables

A confounding variable is a factor that influences both variables being studied.

For example:

For example:

Deprivation
      ↓
Smoking prevalence
      ↓
Stroke risk

If smoking prevalence is not considered, we may overestimate the direct effect of deprivation alone.

Healthcare datasets often contain many confounding variables simultaneously.

This is one reason why interpreting healthcare analytics requires caution.

Why Adjustment Matters

Healthcare analytics often uses adjustment methods to improve fairness and interpretation.

Examples include:

  • age standardisation
  • case-mix adjustment
  • regression modelling
  • risk adjustment

Adjustment helps us ask:

“What would outcomes look like if populations were more comparable?”

For example, after adjusting for:

  • age
  • smoking
  • hypertension
  • diabetes
  • frailty

…the gap in stroke mortality between populations may reduce substantially.

This does not mean inequalities disappear.

It means:

part of the observed relationship was explained by other linked factors.

Adjustment improves understanding — it does not remove complexity.

Looking Deeper

If interested, see:

Appendix — How Adjustment Works

for a more detailed explanation of adjustment approaches used in healthcare analytics.

Correlation Can Still Be Extremely Useful

Importantly:

correlation is not “bad”.

Correlation is often the starting point for:

  • identifying inequalities
  • detecting unwarranted variation
  • prioritising interventions
  • operational investigation
  • pathway redesign

Without correlation analysis:

  • many healthcare inequalities would remain hidden.

The key is understanding that:

correlation should start questions, not end them.

Operational Implications

Misinterpreting correlation as causation can lead to:

  • simplistic conclusions
  • ineffective interventions
  • unrealistic targets
  • inappropriate benchmarking

For example:

If analysis shows:

  • higher ED attendance in deprived areas

…the operational response should not simply be:

“reduce ED demand.”

Instead, systems may need to examine:

  • primary care access
  • prevention services
  • frailty support
  • transport barriers
  • community capacity
  • housing instability
  • proactive care pathways

This moves analytics from:

  • descriptive reporting towards:
  • operational understanding and intervention.

From Correlation to Operational Action

Effective healthcare analytics often follows this progression:

Step Purpose
Correlation Detect a signal
Identify drivers Explore explanations
Adjustment Improve interpretation
Cohort segmentation Identify who is affected
Operational targeting Focus intervention
Pathway redesign Improve outcomes

This is where healthcare analytics becomes most valuable — moving from observation to better intervention design.

Why Healthcare Data Is Particularly Complex

Healthcare outcomes rarely arise from a single cause.

Instead:

  • multiple factors interact
  • pathways overlap
  • operational context matters

This is why healthcare analytics requires:

interpretation, context and operational understanding — not statistical outputs alone

Key Takeaways

  • Correlation describes relationships between variables
  • Correlation alone does not prove causation
  • Correlation should start questions, not end them
  • Healthcare inequalities are usually driven by multiple interacting factors
  • Confounding variables can distort interpretation
  • Adjustment helps improve fairness and understanding
  • Correlation remains extremely valuable for identifying patterns and prioritising investigation
  • Healthcare analytics should support better questions, not simplistic conclusions

Questions Decision-Makers Should Ask

When reviewing correlated healthcare metrics:

  • What factors may sit behind this relationship?
  • Are important confounding variables being missed?
  • Does this represent causation, association, or both?
  • Are populations truly comparable?
  • What operational factors influence these outcomes?
  • Which drivers are actually modifiable?
  • What intervention opportunities exist?
  • What additional analysis is needed before acting?