Module 9 — Evaluation Methods in Practice

How healthcare systems improve confidence in evaluation findings

Module Learning Objective

This module helps explain:

how healthcare systems move beyond simple before-and-after comparisons by using more robust approaches to improve confidence that an intervention genuinely contributed to observed change.

By the end of this module, readers should feel more confident asking:

How strong is the evidence that this intervention worked?

The module focuses on practical healthcare examples, particularly interventions intended to reduce:

  • Emergency Department (ED) attendances
  • emergency admissions
  • avoidable hospital utilisation
  • delayed discharge
  • operational pressure

Rather than assuming healthcare systems can ever prove impact with complete certainty, the module focuses on:

improving confidence in evaluation findings

through more thoughtful and structured approaches.

From Module 7 to Module 8

In Module 7 we introduced an important idea:

improvement alone does not prove impact

Healthcare systems frequently observe change after an intervention.

For example:

ED attendances ↓
Emergency admissions ↓

But an important question remains:

Would this improvement have happened anyway?

This idea introduced:

  • counterfactual thinking
  • confidence in evidence
  • contribution versus attribution
  • the limitations of simple before-and-after comparisons

The question now becomes:

How do healthcare systems improve confidence that an intervention genuinely contributed to observed change?

This introduces practical approaches that are commonly used in healthcare evaluation.

These approaches do not eliminate uncertainty.

Instead, they help decision-makers ask:

How confident should we be?

Why Practical Evaluation Methods Matter

Healthcare systems often need to make decisions quickly.

Programmes are launched:

  • under operational pressure
  • during winter demand
  • amidst workforce shortages
  • alongside other service changes

Decision-makers may need to decide:

  • whether to continue funding a scheme
  • whether to scale it across the system
  • whether to redesign or stop it

Yet waiting for perfect evidence is rarely realistic.

This creates an important challenge:

How do we improve confidence in evaluation findings without waiting years for certainty?

Practical evaluation methods help by:

  • improving comparisons
  • strengthening counterfactual thinking
  • reducing misleading conclusions
  • increasing confidence in observed impact

Importantly:

stronger methods do not guarantee truth.

Instead:

they help reduce the risk of reaching the wrong conclusion.

Randomised Controlled Trials (RCTs)

One of the strongest ways to evaluate whether an intervention caused change is through a:

Randomised Controlled Trial (RCT)

RCTs are often described as the:

gold standard of evaluation

because they help create a stronger counterfactual.

What Is an RCT?

In simple terms:

an RCT compares two groups.

One group receives:

the intervention

The other receives:

usual care (or no intervention)

Participants are assigned randomly.

This is important because randomisation helps reduce bias.

In theory:

the two groups should be broadly similar at the start.

If outcomes later differ:

we can have greater confidence that:

the intervention contributed to the change

rather than:

  • chance
  • selection bias
  • population differences
  • other confounding factors

A Practical Healthcare Example

Imagine an Integrated Care Board introduces a frailty intervention designed to reduce emergency admissions.

Rather than implementing the service everywhere immediately:

several Primary Care Networks (PCNs) are selected to participate.

Patients are randomly allocated.

Group A

Receives:

  • proactive frailty review
  • medication optimisation
  • MDT support
  • community monitoring

Group B

Receives:

usual care

After twelve months:

Measure Intervention Group Usual Care
ED attendances 900 1,100
Emergency admissions 320 420

Question:

Did the intervention contribute to improvement?

Because patients were randomly allocated, we can generally have:

greater confidence that observed differences are more likely to reflect the intervention rather than underlying differences between patients.

Why Randomisation Matters

Without randomisation:

services may unintentionally select patients who are:

  • easier to help
  • more engaged
  • less clinically complex
  • more motivated

This creates a problem known as:

selection bias

For example:

a frailty service might prioritise patients who are already likely to improve.

The intervention may then appear more successful than it truly is.

Randomisation helps reduce this risk.

The key idea is simple:

RCTs attempt to improve fairness in comparison by reducing differences between groups before outcomes are compared.

random allocation improves confidence that differences in outcomes are more likely to reflect the intervention rather than underlying differences between patients.

Why RCTs Are Difficult in Real Healthcare Systems

Despite their strengths:

RCTs are often difficult to implement in real healthcare systems.

RCTs can provide strong evidence.

However:

strong evidence is not always operationally practical

Healthcare environments are messy.

Operational realities include:

  • workforce shortages
  • service redesign during implementation
  • changing pathways
  • political and operational pressure to scale interventions quickly
  • ethical concerns about withholding care

For example:

if early evidence suggests a service is beneficial:

leaders may feel uncomfortable withholding support from some patients.

Similarly:

healthcare systems often need to move quickly.

Waiting several years for a perfect evaluation may not feel realistic when:

  • Emergency Departments are under pressure
  • emergency admissions are rising
  • operational performance is deteriorating

As a result:

many healthcare evaluations rely on:

pragmatic approaches

that improve confidence without requiring a full RCT.

Key Message

RCTs are powerful because they help create stronger comparisons between intervention and non-intervention groups.

However:

strong evidence is not always operationally practical evidence

Healthcare systems often need approaches that balance:

rigour

with

operational reality

Part 1 Reflection

Before moving on, ask:

If a scheme appears successful, how confident are we that the groups being compared were genuinely comparable?

In the next section, we explore another important question:

How much evidence is enough to trust the findings?

This introduces:

sample size and statistical power

Sample Size, Statistical Power and Confidence in Findings

In 6, we introduced an important question:

How confident are we that an intervention genuinely contributed to observed change?

Even when groups are fairly compared, another important challenge remains:

Do we have enough evidence to trust what we are seeing?

This is where:

sample size and statistical power

become important.

Fortunately, the core idea is much simpler than the terminology suggests.

Why Small Numbers Can Mislead

Healthcare systems often evaluate schemes using:

  • small pilots
  • short implementation periods
  • highly selected patient groups
  • limited operational data

For example:

A frailty intervention is piloted for:

20 patients

After three months:

Measure Before After
ED attendances 42 28
Emergency admissions 17 10

At first glance:

the scheme appears successful.

But an important question remains:

How confident should we be that this change reflects genuine impact rather than normal fluctuation?

Small numbers naturally fluctuate.

A handful of patients can disproportionately influence results.

For example:

if two highly complex patients experience:

  • hospitalisation
  • bereavement
  • worsening frailty
  • medication changes

outcomes may change substantially.

Similarly:

if several patients improve naturally:

results may look highly positive.

This means:

Small pilots may produce convincing-looking results that are less reliable than they first appear.

What Do We Mean by Sample Size?

Sample size simply refers to:

how many people, events or observations are included in an evaluation

Examples include:

  • number of patients enrolled in a scheme
  • number of Emergency Department attendances analysed
  • number of GP practices or PCNs involved
  • number of months of activity included

In general:

larger samples tend to produce more stable and reliable findings

because random fluctuation becomes less influential.

This does not mean:

larger samples automatically guarantee better evidence

Poor evaluation design can still produce misleading conclusions.

However:

small numbers deserve greater caution

Statistical Power (In Plain English)

Statistical power sounds technical.

But the underlying question is straightforward:

If a real improvement exists, how likely are we to detect it?

Low-powered evaluations may fail to detect genuine effects.

For example:

a frailty scheme may genuinely reduce emergency admissions by:

5%

But if:

  • only a small number of patients are included
  • follow-up time is short
  • variation is high

the evaluation may struggle to detect the change.

Decision-makers may incorrectly conclude:

The intervention didn’t work

when the evaluation simply lacked enough information.

In practice:

small evaluations are more likely to miss real effects — or overreact to random fluctuation.

A Practical Healthcare Example

Imagine two evaluations of the same frailty pathway.

Evaluation A

20 patients
2 months follow-up

Observed result:

Emergency admissions ↓

Question:

Was this meaningful improvement — or random fluctuation?

Confidence remains limited.

Evaluation B

2,000 patients
12 months follow-up
Comparator group included

Observed result:

Emergency admissions ↓

Question:

How confident are we now?

Confidence is greater because:

  • more patients were observed
  • longer time period analysed
  • natural fluctuation matters less
  • comparison is stronger

This does not guarantee truth.

But it increases confidence.

The image below illustrates why small evaluations often produce more unstable and uncertain findings — even when results initially look promising.

Why This Matters for Healthcare Decision-Making

Healthcare systems often pilot schemes quickly.

Leaders may face pressure to decide:

  • continue funding
  • scale across the system
  • redesign
  • stop

before strong evidence exists.

This creates an important tension:

speed of learning

versus

confidence in conclusions

The question is rarely:

Can we be completely certain?

Instead:

How confident should we be before acting?

Regression to the Mean

Even when evaluations include enough observations, another important challenge remains:

could improvement have happened naturally anyway?

This brings us to:

Regression to the Mean

The name sounds complicated.

The idea is not.

It describes a simple tendency:

extreme outcomes often become less extreme over time

even without intervention.

This matters enormously in healthcare.

Because interventions are often targeted at:

  • high-risk patients
  • high-intensity users
  • people with repeated admissions
  • people with severe frailty

In other words:

the very people already experiencing unusually extreme outcomes

A Practical Healthcare Example

Imagine a healthcare system identifies:

20 patients

with very high Emergency Department attendance.

Each attended:

10–15 times

during the previous year.

A targeted intervention is introduced.

For example:

  • care coordination
  • frailty MDT support
  • proactive monitoring
  • social prescribing
  • medication review

One year later:

ED attendance falls.

Question:

Did the intervention work?

Maybe.

But another possibility exists.

Patients selected because they experienced unusually extreme healthcare utilisation may naturally move closer to average levels over time.

Some people improve naturally because:

  • acute illness resolves
  • social circumstances stabilise
  • medication changes occur
  • temporary crises pass

This means:

some improvement may occur even if the intervention had little or no effect

Why Regression to the Mean Matters

Without recognising regression to the mean:

healthcare systems may mistakenly conclude:

The scheme worked

when some improvement might have happened anyway.

This is particularly important when evaluating:

  • high-intensity ED attenders
  • frequent emergency admissions
  • severe frailty cohorts
  • complex multimorbidity pathways

The higher the starting level of risk:

the greater the risk of mistaking natural fluctuation for intervention impact

The image below illustrates an important evaluation risk:

improvement does not always mean intervention impact.

Why This Matters for Healthcare Decision-Making

Healthcare systems frequently target interventions at patients with the highest levels of need.

This makes regression to the mean particularly important.

Without recognising it, decision-makers may:

  • overestimate intervention impact
  • scale ineffective schemes
  • misallocate scarce resources

A useful question becomes:

Would some improvement have happened anyway?

Key Message

Small pilots and extreme patient groups can easily mislead evaluation.

Healthcare systems should therefore ask:

Are we observing genuine intervention impact — or natural fluctuation?

Good evaluation becomes stronger when we ask:

Do we have enough evidence?

and

Could improvement have happened anyway?

Part 2 — Confidence, Small Numbers & Natural Fluctuation

Before moving on, ask:

If outcomes improved, how confident are we that this reflects intervention impact rather than small numbers or natural fluctuation?

In the next section, we explore practical ways healthcare systems strengthen confidence in evaluation findings.

Part 3 — Strengthening Evaluation in Real-World Healthcare

In Part 1 and Part 2 of this , we explored:

  • fair comparison through Randomised Controlled Trials (RCTs)
  • why small numbers can mislead
  • statistical power
  • Regression to the Mean

Yet another challenge remains.

Healthcare systems rarely operate in perfect conditions.

Leaders often need to evaluate schemes:

  • quickly
  • under operational pressure
  • without randomisation
  • while multiple changes happen simultaneously

In reality:

RCTs are often difficult, expensive or impractical

This creates an important question:

How can healthcare systems strengthen evaluation when RCTs are unrealistic?

Two practical approaches commonly used in healthcare evaluation are:

Interrupted Time Series (ITS)

and

Difference-in-Differences (DiD)

These approaches do not eliminate uncertainty.

Instead:

they improve confidence in findings by creating stronger comparisons.

Interrupted Time Series (ITS)

One challenge with simple before-and-after evaluation is this:

healthcare systems naturally fluctuate over time.

Emergency admissions rise and fall.

Waiting lists change.

Operational pressure changes.

Winter pressures come and go.

This creates an important risk:

improvement may already have been happening before an intervention started

Interrupted Time Series (ITS) helps address this problem.

What Is Interrupted Time Series?

Rather than asking:

Did outcomes improve?

ITS asks:

Did the pattern of change meaningfully alter after the intervention?

This distinction matters.

Instead of simply comparing:

before
vs
after

ITS asks:

Was the trajectory interrupted?

In simple terms:

ITS looks at:

  • trends before intervention
  • trends after intervention
  • whether the pattern changed meaningfully

For example:

Did emergency admissions:

  • suddenly reduce?
  • begin falling faster?
  • stop increasing?
  • improve in a sustained way?

A Practical Healthcare Example

Imagine an Integrated Care Board introduces a frailty pathway to reduce emergency admissions.

Instead of comparing:

before vs after

we examine:

monthly emergency admissions over time

Suppose the intervention begins in:

Month 8

Question:

Did emergency admissions meaningfully change after implementation?

Three possibilities might emerge:

Scenario 1 — No Meaningful Change

Admissions were already falling.

The trend simply continues.

Question:

Did the intervention add anything?

Confidence in intervention impact remains limited.

Scenario 2 — Sudden Change

Admissions drop sharply immediately after implementation.

Question:

Could this represent intervention impact?

Confidence increases.

Scenario 3 — Sustained Improvement

Admissions begin falling faster after implementation and remain lower over time.

Question:

Does this suggest a meaningful shift in trajectory?

Confidence increases further.

The image below illustrates the key question ITS asks:

Did the trajectory meaningfully change after intervention?

Interrupted Time Series helps distinguish:

normal change over time

from

potential intervention impact

Why ITS Matters for Healthcare Decision-Making

Healthcare systems often mistake:

normal trends

for

intervention impact

For example:

Emergency admissions may naturally improve after winter.

Operational performance may recover after temporary disruption.

Waiting lists may already have been stabilising.

Without considering trends:

decision-makers risk concluding:

The scheme worked

when outcomes may have improved anyway.

A useful question becomes:

Did outcomes improve — or did the trajectory change?

Using Interrupted Time Series in Practice

Interrupted Time Series is one of the strongest quasi-experimental designs available when randomisation is not possible.

However, choosing an appropriate ITS approach requires careful consideration of both the data and the evaluation question.

Before fitting an ITS model, analysts should ask:

  • Do we have enough observations before and after the intervention?
  • Is there a clearly defined intervention date?
  • Were other important changes introduced at the same time?
  • Is there evidence of seasonality (for example, winter pressures)?
  • Are outcomes measured consistently over time?
  • Has the underlying trend been stable before the intervention?
  • Could changes simply reflect natural variation rather than intervention impact?

These questions help determine whether Interrupted Time Series is likely to provide a meaningful estimate of intervention impact.

TipChoosing the Right Evaluation Design

Good evaluation begins with the evaluation question, not the statistical model.

The aim is not simply to fit an Interrupted Time Series model.

The aim is to select a design that provides the most credible estimate of what would probably have happened had the intervention not occurred.

Evaluating Multiple Organisations

Healthcare evaluations often involve several hospitals, Primary Care Networks or localities.

This introduces additional complexity.

For example:

  • Different organisations may begin with different baseline performance.
  • Improvement may already be occurring at different rates.
  • Interventions may be implemented slightly differently.
  • Local operational pressures may vary considerably.

There is rarely a single correct analytical approach.

Depending on the evaluation question and available data, analysts may choose to:

  • analyse each organisation separately
  • fit a single model that accounts for organisational differences
  • combine findings across organisations using evidence synthesis techniques

The choice depends on the evaluation question, study design and available data.

The important principle is not selecting the most complex model, but choosing the approach that provides the most credible answer to the evaluation question.

NoteLooking Ahead

Real-world healthcare improvement rarely involves a single intervention delivered in a single organisation.

Later in this series, Module 15: Evaluating Change in Complex Healthcare Systems explores more advanced evaluation challenges, including:

  • multiple interventions introduced at the same time
  • phased implementation across organisations
  • multi-site evaluations
  • mixed-method evaluation
  • attribution versus contribution

These situations require careful evaluation design as well as appropriate statistical methods.

Difference-in-Differences (DiD)

Another practical challenge in healthcare evaluation is this:

Compared with what?

Even when outcomes improve:

another question remains:

Would similar improvement have happened elsewhere anyway?

Difference-in-Differences (DiD) helps strengthen confidence by introducing:

a comparator group

What Is Difference-in-Differences?

Rather than asking:

Did outcomes improve?

DiD asks:

Did outcomes improve more where the intervention happened?

In simple terms:

we compare:

change
vs
change

rather than:

before
vs
after

The idea is simple:

If two similar groups experience similar pressures —

but only one receives the intervention —

then differences in improvement may provide stronger evidence of impact.

A Practical Healthcare Example

Imagine:

PCN A

Introduces a frailty pathway.

PCN B

Does not.

Both experience:

  • winter pressure
  • workforce shortages
  • rising demand

After twelve months:

PCN A

Emergency admissions ↓ 15%

PCN B

Emergency admissions ↓ 5%

Question:

Did the intervention contribute to additional improvement?

The key question becomes:

Was improvement greater where the intervention occurred?

This provides stronger evidence than:

before-and-after comparison alone

The image below illustrates the core idea behind Difference-in-Differences:

good evaluation compares change — not just outcomes

Why DiD Matters for Healthcare Decision-Making

Healthcare systems frequently implement interventions:

  • gradually
  • in some places before others
  • under different operational conditions

This makes DiD especially useful when:

  • RCTs are unrealistic
  • phased roll-out exists
  • comparator organisations or places are available

Without comparators:

decision-makers risk attributing:

system-wide improvement

to

local intervention impact

A useful question becomes:

Would similar improvement have happened elsewhere anyway?

Key Message

Good evaluation becomes stronger when healthcare systems ask:

Did trajectories change?

and

Compared with what?

rather than simply asking:

Did outcomes improve?

Pragmatic evaluation methods such as:

Interrupted Time Series (ITS)

and

Difference-in-Differences (DiD)

help strengthen confidence when Randomised Controlled Trials are not practical.

Part 3 Reflection

Before acting on evaluation findings, ask:

Could this improvement reflect an existing trend — or would similar change have happened elsewhere anyway?

Good evaluation moves beyond:

Did outcomes improve?

to ask:

How confident are we that the intervention genuinely contributed to change?

Matching the Evaluation Method to the Question

One of the most important principles in evaluation is that the evaluation question should determine the evaluation method.

There is rarely a single “best” evaluation method.

Instead, different methods answer different questions.

Underlying them all is one fundamental idea:

What would probably have happened if the intervention had not taken place?

This is known as the counterfactual.

Different evaluation methods estimate the counterfactual in different ways.

Evaluation Question How is the Counterfactual Considered? Methods Commonly Used
Did it work? Compare outcomes with what would probably have happened without the intervention. Randomised Controlled Trial (RCT), Interrupted Time Series (ITS), Difference-in-Differences (DiD), Controlled ITS
How was it implemented? Understand how the intervention was delivered and whether implementation influenced outcomes. Process evaluation, implementation evaluation, qualitative interviews
Why did it work (or not)? Explore the mechanisms that contributed to success or failure. Qualitative research, Realist Evaluation
Did it work consistently across organisations? Compare patterns of change across multiple organisations while accounting for differences in starting position, trends and implementation. Multi-site evaluation, mixed-effects models, meta-analysis
Was it worth it? Compare the benefits achieved against the resources invested. Health economics, Cost-effectiveness analysis, cost-benefit analysis, cost-utility analysis
Should we scale it? Combine evidence on effectiveness, implementation, value and context to judge wider adoption. Mixed-method evaluation, evidence synthesis, implementation evaluation
TipKey Principle

Good evaluation starts with a clear evaluation question, not a preferred method.

Strong evaluations often combine quantitative and qualitative approaches to understand:

  • whether an intervention worked
  • how it was implemented
  • why it worked (or did not work)
  • whether it represented good value for money
  • whether it should be adopted more widely

Although different methods estimate the counterfactual in different ways, they all aim to increase confidence that observed changes genuinely resulted from the intervention rather than from chance, natural variation or external influences.

The stronger the evaluation design, the greater our confidence in the conclusions—but no evaluation can eliminate uncertainty completely.

Key Takeaways

  • improvement alone does not prove impact
  • fair comparison matters
  • small numbers can mislead
  • some improvement may happen naturally
  • trajectories matter
  • comparison groups strengthen confidence
  • the evaluation question should determine the evaluation method

Good evaluation asks:

How confident are we that the intervention genuinely contributed to change?

Questions Decision-Makers Should Ask

  • Compared with what?
  • Did the trajectory change?
  • Could improvement have happened naturally?
  • Was the evaluation large enough to be informative?
  • How strong is our confidence in attribution?

Choosing the Right Evaluation Method

Further Reading

Readers who would like to explore evaluation methods in greater depth may find the following resources useful: