Module 9 — Evaluation Methods in Practice
How healthcare systems improve confidence in evaluation findings
Module Learning Objective
This module helps explain:
how healthcare systems move beyond simple before-and-after comparisons by using more robust approaches to improve confidence that an intervention genuinely contributed to observed change.
By the end of this module, readers should feel more confident asking:
How strong is the evidence that this intervention worked?
The module focuses on practical healthcare examples, particularly interventions intended to reduce:
- Emergency Department (ED) attendances
- emergency admissions
- avoidable hospital utilisation
- delayed discharge
- operational pressure
Rather than assuming healthcare systems can ever prove impact with complete certainty, the module focuses on:
improving confidence in evaluation findings
through more thoughtful and structured approaches.
From Module 7 to Module 8
In Module 7 we introduced an important idea:
improvement alone does not prove impact
Healthcare systems frequently observe change after an intervention.
For example:
ED attendances ↓
Emergency admissions ↓
But an important question remains:
Would this improvement have happened anyway?
This idea introduced:
- counterfactual thinking
- confidence in evidence
- contribution versus attribution
- the limitations of simple before-and-after comparisons
The question now becomes:
How do healthcare systems improve confidence that an intervention genuinely contributed to observed change?
This introduces practical approaches that are commonly used in healthcare evaluation.
These approaches do not eliminate uncertainty.
Instead, they help decision-makers ask:
How confident should we be?
Why Practical Evaluation Methods Matter
Healthcare systems often need to make decisions quickly.
Programmes are launched:
- under operational pressure
- during winter demand
- amidst workforce shortages
- alongside other service changes
Decision-makers may need to decide:
- whether to continue funding a scheme
- whether to scale it across the system
- whether to redesign or stop it
Yet waiting for perfect evidence is rarely realistic.
This creates an important challenge:
How do we improve confidence in evaluation findings without waiting years for certainty?
Practical evaluation methods help by:
- improving comparisons
- strengthening counterfactual thinking
- reducing misleading conclusions
- increasing confidence in observed impact
Importantly:
stronger methods do not guarantee truth.
Instead:
they help reduce the risk of reaching the wrong conclusion.
Randomised Controlled Trials (RCTs)
One of the strongest ways to evaluate whether an intervention caused change is through a:
Randomised Controlled Trial (RCT)
RCTs are often described as the:
gold standard of evaluation
because they help create a stronger counterfactual.
What Is an RCT?
In simple terms:
an RCT compares two groups.
One group receives:
the intervention
The other receives:
usual care (or no intervention)
Participants are assigned randomly.
This is important because randomisation helps reduce bias.
In theory:
the two groups should be broadly similar at the start.
If outcomes later differ:
we can have greater confidence that:
the intervention contributed to the change
rather than:
- chance
- selection bias
- population differences
- other confounding factors
A Practical Healthcare Example
Imagine an Integrated Care Board introduces a frailty intervention designed to reduce emergency admissions.
Rather than implementing the service everywhere immediately:
several Primary Care Networks (PCNs) are selected to participate.
Patients are randomly allocated.
Group A
Receives:
- proactive frailty review
- medication optimisation
- MDT support
- community monitoring
Group B
Receives:
usual care
After twelve months:
| Measure | Intervention Group | Usual Care |
|---|---|---|
| ED attendances | 900 | 1,100 |
| Emergency admissions | 320 | 420 |
Question:
Did the intervention contribute to improvement?
Because patients were randomly allocated, we can generally have:
greater confidence that observed differences are more likely to reflect the intervention rather than underlying differences between patients.
Why Randomisation Matters
Without randomisation:
services may unintentionally select patients who are:
- easier to help
- more engaged
- less clinically complex
- more motivated
This creates a problem known as:
selection bias
For example:
a frailty service might prioritise patients who are already likely to improve.
The intervention may then appear more successful than it truly is.
Randomisation helps reduce this risk.

The key idea is simple:
RCTs attempt to improve fairness in comparison by reducing differences between groups before outcomes are compared.
random allocation improves confidence that differences in outcomes are more likely to reflect the intervention rather than underlying differences between patients.
Why RCTs Are Difficult in Real Healthcare Systems
Despite their strengths:
RCTs are often difficult to implement in real healthcare systems.
RCTs can provide strong evidence.
However:
strong evidence is not always operationally practical
Healthcare environments are messy.
Operational realities include:
- workforce shortages
- service redesign during implementation
- changing pathways
- political and operational pressure to scale interventions quickly
- ethical concerns about withholding care
For example:
if early evidence suggests a service is beneficial:
leaders may feel uncomfortable withholding support from some patients.
Similarly:
healthcare systems often need to move quickly.
Waiting several years for a perfect evaluation may not feel realistic when:
- Emergency Departments are under pressure
- emergency admissions are rising
- operational performance is deteriorating
As a result:
many healthcare evaluations rely on:
pragmatic approaches
that improve confidence without requiring a full RCT.
Key Message
RCTs are powerful because they help create stronger comparisons between intervention and non-intervention groups.
However:
strong evidence is not always operationally practical evidence
Healthcare systems often need approaches that balance:
rigour
with
operational reality
Part 1 Reflection
Before moving on, ask:
If a scheme appears successful, how confident are we that the groups being compared were genuinely comparable?
In the next section, we explore another important question:
How much evidence is enough to trust the findings?
This introduces:
sample size and statistical power
Sample Size, Statistical Power and Confidence in Findings
In 6, we introduced an important question:
How confident are we that an intervention genuinely contributed to observed change?
Even when groups are fairly compared, another important challenge remains:
Do we have enough evidence to trust what we are seeing?
This is where:
sample size and statistical power
become important.
Fortunately, the core idea is much simpler than the terminology suggests.
Why Small Numbers Can Mislead
Healthcare systems often evaluate schemes using:
- small pilots
- short implementation periods
- highly selected patient groups
- limited operational data
For example:
A frailty intervention is piloted for:
20 patients
After three months:
| Measure | Before | After |
|---|---|---|
| ED attendances | 42 | 28 |
| Emergency admissions | 17 | 10 |
At first glance:
the scheme appears successful.
But an important question remains:
How confident should we be that this change reflects genuine impact rather than normal fluctuation?
Small numbers naturally fluctuate.
A handful of patients can disproportionately influence results.
For example:
if two highly complex patients experience:
- hospitalisation
- bereavement
- worsening frailty
- medication changes
outcomes may change substantially.
Similarly:
if several patients improve naturally:
results may look highly positive.
This means:
Small pilots may produce convincing-looking results that are less reliable than they first appear.
What Do We Mean by Sample Size?
Sample size simply refers to:
how many people, events or observations are included in an evaluation
Examples include:
- number of patients enrolled in a scheme
- number of Emergency Department attendances analysed
- number of GP practices or PCNs involved
- number of months of activity included
In general:
larger samples tend to produce more stable and reliable findings
because random fluctuation becomes less influential.
This does not mean:
larger samples automatically guarantee better evidence
Poor evaluation design can still produce misleading conclusions.
However:
small numbers deserve greater caution
Statistical Power (In Plain English)
Statistical power sounds technical.
But the underlying question is straightforward:
If a real improvement exists, how likely are we to detect it?
Low-powered evaluations may fail to detect genuine effects.
For example:
a frailty scheme may genuinely reduce emergency admissions by:
5%
But if:
- only a small number of patients are included
- follow-up time is short
- variation is high
the evaluation may struggle to detect the change.
Decision-makers may incorrectly conclude:
The intervention didn’t work
when the evaluation simply lacked enough information.
In practice:
small evaluations are more likely to miss real effects — or overreact to random fluctuation.
A Practical Healthcare Example
Imagine two evaluations of the same frailty pathway.
Evaluation A
20 patients
2 months follow-up
Observed result:
Emergency admissions ↓
Question:
Was this meaningful improvement — or random fluctuation?
Confidence remains limited.
Evaluation B
2,000 patients
12 months follow-up
Comparator group included
Observed result:
Emergency admissions ↓
Question:
How confident are we now?
Confidence is greater because:
- more patients were observed
- longer time period analysed
- natural fluctuation matters less
- comparison is stronger
This does not guarantee truth.
But it increases confidence.
The image below illustrates why small evaluations often produce more unstable and uncertain findings — even when results initially look promising.

Why This Matters for Healthcare Decision-Making
Healthcare systems often pilot schemes quickly.
Leaders may face pressure to decide:
- continue funding
- scale across the system
- redesign
- stop
before strong evidence exists.
This creates an important tension:
speed of learning
versus
confidence in conclusions
The question is rarely:
Can we be completely certain?
Instead:
How confident should we be before acting?
Regression to the Mean
Even when evaluations include enough observations, another important challenge remains:
could improvement have happened naturally anyway?
This brings us to:
Regression to the Mean
The name sounds complicated.
The idea is not.
It describes a simple tendency:
extreme outcomes often become less extreme over time
even without intervention.
This matters enormously in healthcare.
Because interventions are often targeted at:
- high-risk patients
- high-intensity users
- people with repeated admissions
- people with severe frailty
In other words:
the very people already experiencing unusually extreme outcomes
A Practical Healthcare Example
Imagine a healthcare system identifies:
20 patients
with very high Emergency Department attendance.
Each attended:
10–15 times
during the previous year.
A targeted intervention is introduced.
For example:
- care coordination
- frailty MDT support
- proactive monitoring
- social prescribing
- medication review
One year later:
ED attendance falls.
Question:
Did the intervention work?
Maybe.
But another possibility exists.
Patients selected because they experienced unusually extreme healthcare utilisation may naturally move closer to average levels over time.
Some people improve naturally because:
- acute illness resolves
- social circumstances stabilise
- medication changes occur
- temporary crises pass
This means:
some improvement may occur even if the intervention had little or no effect
Why Regression to the Mean Matters
Without recognising regression to the mean:
healthcare systems may mistakenly conclude:
The scheme worked
when some improvement might have happened anyway.
This is particularly important when evaluating:
- high-intensity ED attenders
- frequent emergency admissions
- severe frailty cohorts
- complex multimorbidity pathways
The higher the starting level of risk:
the greater the risk of mistaking natural fluctuation for intervention impact
The image below illustrates an important evaluation risk:
improvement does not always mean intervention impact.

Why This Matters for Healthcare Decision-Making
Healthcare systems frequently target interventions at patients with the highest levels of need.
This makes regression to the mean particularly important.
Without recognising it, decision-makers may:
- overestimate intervention impact
- scale ineffective schemes
- misallocate scarce resources
A useful question becomes:
Would some improvement have happened anyway?
Key Message
Small pilots and extreme patient groups can easily mislead evaluation.
Healthcare systems should therefore ask:
Are we observing genuine intervention impact — or natural fluctuation?
Good evaluation becomes stronger when we ask:
Do we have enough evidence?
and
Could improvement have happened anyway?
Part 2 — Confidence, Small Numbers & Natural Fluctuation
Before moving on, ask:
If outcomes improved, how confident are we that this reflects intervention impact rather than small numbers or natural fluctuation?
In the next section, we explore practical ways healthcare systems strengthen confidence in evaluation findings.
Part 3 — Strengthening Evaluation in Real-World Healthcare
In Part 1 and Part 2 of this , we explored:
- fair comparison through Randomised Controlled Trials (RCTs)
- why small numbers can mislead
- statistical power
- Regression to the Mean
Yet another challenge remains.
Healthcare systems rarely operate in perfect conditions.
Leaders often need to evaluate schemes:
- quickly
- under operational pressure
- without randomisation
- while multiple changes happen simultaneously
In reality:
RCTs are often difficult, expensive or impractical
This creates an important question:
How can healthcare systems strengthen evaluation when RCTs are unrealistic?
Two practical approaches commonly used in healthcare evaluation are:
Interrupted Time Series (ITS)
and
Difference-in-Differences (DiD)
These approaches do not eliminate uncertainty.
Instead:
they improve confidence in findings by creating stronger comparisons.
Interrupted Time Series (ITS)
One challenge with simple before-and-after evaluation is this:
healthcare systems naturally fluctuate over time.
Emergency admissions rise and fall.
Waiting lists change.
Operational pressure changes.
Winter pressures come and go.
This creates an important risk:
improvement may already have been happening before an intervention started
Interrupted Time Series (ITS) helps address this problem.
What Is Interrupted Time Series?
Rather than asking:
Did outcomes improve?
ITS asks:
Did the pattern of change meaningfully alter after the intervention?
This distinction matters.
Instead of simply comparing:
before
vs
after
ITS asks:
Was the trajectory interrupted?
In simple terms:
ITS looks at:
- trends before intervention
- trends after intervention
- whether the pattern changed meaningfully
For example:
Did emergency admissions:
- suddenly reduce?
- begin falling faster?
- stop increasing?
- improve in a sustained way?
A Practical Healthcare Example
Imagine an Integrated Care Board introduces a frailty pathway to reduce emergency admissions.
Instead of comparing:
before vs after
we examine:
monthly emergency admissions over time
Suppose the intervention begins in:
Month 8
Question:
Did emergency admissions meaningfully change after implementation?
Three possibilities might emerge:
Scenario 1 — No Meaningful Change
Admissions were already falling.
The trend simply continues.
Question:
Did the intervention add anything?
Confidence in intervention impact remains limited.
Scenario 2 — Sudden Change
Admissions drop sharply immediately after implementation.
Question:
Could this represent intervention impact?
Confidence increases.
Scenario 3 — Sustained Improvement
Admissions begin falling faster after implementation and remain lower over time.
Question:
Does this suggest a meaningful shift in trajectory?
Confidence increases further.
The image below illustrates the key question ITS asks:
Did the trajectory meaningfully change after intervention?
Interrupted Time Series helps distinguish:
normal change over time
from
potential intervention impact

Why ITS Matters for Healthcare Decision-Making
Healthcare systems often mistake:
normal trends
for
intervention impact
For example:
Emergency admissions may naturally improve after winter.
Operational performance may recover after temporary disruption.
Waiting lists may already have been stabilising.
Without considering trends:
decision-makers risk concluding:
The scheme worked
when outcomes may have improved anyway.
A useful question becomes:
Did outcomes improve — or did the trajectory change?
Using Interrupted Time Series in Practice
Interrupted Time Series is one of the strongest quasi-experimental designs available when randomisation is not possible.
However, choosing an appropriate ITS approach requires careful consideration of both the data and the evaluation question.
Before fitting an ITS model, analysts should ask:
- Do we have enough observations before and after the intervention?
- Is there a clearly defined intervention date?
- Were other important changes introduced at the same time?
- Is there evidence of seasonality (for example, winter pressures)?
- Are outcomes measured consistently over time?
- Has the underlying trend been stable before the intervention?
- Could changes simply reflect natural variation rather than intervention impact?
These questions help determine whether Interrupted Time Series is likely to provide a meaningful estimate of intervention impact.
Good evaluation begins with the evaluation question, not the statistical model.
The aim is not simply to fit an Interrupted Time Series model.
The aim is to select a design that provides the most credible estimate of what would probably have happened had the intervention not occurred.
Evaluating Multiple Organisations
Healthcare evaluations often involve several hospitals, Primary Care Networks or localities.
This introduces additional complexity.
For example:
- Different organisations may begin with different baseline performance.
- Improvement may already be occurring at different rates.
- Interventions may be implemented slightly differently.
- Local operational pressures may vary considerably.
There is rarely a single correct analytical approach.
Depending on the evaluation question and available data, analysts may choose to:
- analyse each organisation separately
- fit a single model that accounts for organisational differences
- combine findings across organisations using evidence synthesis techniques
The choice depends on the evaluation question, study design and available data.
The important principle is not selecting the most complex model, but choosing the approach that provides the most credible answer to the evaluation question.
Real-world healthcare improvement rarely involves a single intervention delivered in a single organisation.
Later in this series, Module 15: Evaluating Change in Complex Healthcare Systems explores more advanced evaluation challenges, including:
- multiple interventions introduced at the same time
- phased implementation across organisations
- multi-site evaluations
- mixed-method evaluation
- attribution versus contribution
These situations require careful evaluation design as well as appropriate statistical methods.
Difference-in-Differences (DiD)
Another practical challenge in healthcare evaluation is this:
Compared with what?
Even when outcomes improve:
another question remains:
Would similar improvement have happened elsewhere anyway?
Difference-in-Differences (DiD) helps strengthen confidence by introducing:
a comparator group
What Is Difference-in-Differences?
Rather than asking:
Did outcomes improve?
DiD asks:
Did outcomes improve more where the intervention happened?
In simple terms:
we compare:
change
vs
change
rather than:
before
vs
after
The idea is simple:
If two similar groups experience similar pressures —
but only one receives the intervention —
then differences in improvement may provide stronger evidence of impact.
A Practical Healthcare Example
Imagine:
PCN A
Introduces a frailty pathway.
PCN B
Does not.
Both experience:
- winter pressure
- workforce shortages
- rising demand
After twelve months:
PCN A
Emergency admissions ↓ 15%
PCN B
Emergency admissions ↓ 5%
Question:
Did the intervention contribute to additional improvement?
The key question becomes:
Was improvement greater where the intervention occurred?
This provides stronger evidence than:
before-and-after comparison alone
The image below illustrates the core idea behind Difference-in-Differences:
good evaluation compares change — not just outcomes

Why DiD Matters for Healthcare Decision-Making
Healthcare systems frequently implement interventions:
- gradually
- in some places before others
- under different operational conditions
This makes DiD especially useful when:
- RCTs are unrealistic
- phased roll-out exists
- comparator organisations or places are available
Without comparators:
decision-makers risk attributing:
system-wide improvement
to
local intervention impact
A useful question becomes:
Would similar improvement have happened elsewhere anyway?
Key Message
Good evaluation becomes stronger when healthcare systems ask:
Did trajectories change?
and
Compared with what?
rather than simply asking:
Did outcomes improve?
Pragmatic evaluation methods such as:
Interrupted Time Series (ITS)
and
Difference-in-Differences (DiD)
help strengthen confidence when Randomised Controlled Trials are not practical.
Part 3 Reflection
Before acting on evaluation findings, ask:
Could this improvement reflect an existing trend — or would similar change have happened elsewhere anyway?
Good evaluation moves beyond:
Did outcomes improve?
to ask:
How confident are we that the intervention genuinely contributed to change?
Matching the Evaluation Method to the Question
One of the most important principles in evaluation is that the evaluation question should determine the evaluation method.
There is rarely a single “best” evaluation method.
Instead, different methods answer different questions.
Underlying them all is one fundamental idea:
What would probably have happened if the intervention had not taken place?
This is known as the counterfactual.
Different evaluation methods estimate the counterfactual in different ways.
| Evaluation Question | How is the Counterfactual Considered? | Methods Commonly Used |
|---|---|---|
| Did it work? | Compare outcomes with what would probably have happened without the intervention. | Randomised Controlled Trial (RCT), Interrupted Time Series (ITS), Difference-in-Differences (DiD), Controlled ITS |
| How was it implemented? | Understand how the intervention was delivered and whether implementation influenced outcomes. | Process evaluation, implementation evaluation, qualitative interviews |
| Why did it work (or not)? | Explore the mechanisms that contributed to success or failure. | Qualitative research, Realist Evaluation |
| Did it work consistently across organisations? | Compare patterns of change across multiple organisations while accounting for differences in starting position, trends and implementation. | Multi-site evaluation, mixed-effects models, meta-analysis |
| Was it worth it? | Compare the benefits achieved against the resources invested. | Health economics, Cost-effectiveness analysis, cost-benefit analysis, cost-utility analysis |
| Should we scale it? | Combine evidence on effectiveness, implementation, value and context to judge wider adoption. | Mixed-method evaluation, evidence synthesis, implementation evaluation |
Good evaluation starts with a clear evaluation question, not a preferred method.
Strong evaluations often combine quantitative and qualitative approaches to understand:
- whether an intervention worked
- how it was implemented
- why it worked (or did not work)
- whether it represented good value for money
- whether it should be adopted more widely
Although different methods estimate the counterfactual in different ways, they all aim to increase confidence that observed changes genuinely resulted from the intervention rather than from chance, natural variation or external influences.
The stronger the evaluation design, the greater our confidence in the conclusions—but no evaluation can eliminate uncertainty completely.
Key Takeaways
- improvement alone does not prove impact
- fair comparison matters
- small numbers can mislead
- some improvement may happen naturally
- trajectories matter
- comparison groups strengthen confidence
- the evaluation question should determine the evaluation method
Good evaluation asks:
How confident are we that the intervention genuinely contributed to change?
Questions Decision-Makers Should Ask
- Compared with what?
- Did the trajectory change?
- Could improvement have happened naturally?
- Was the evaluation large enough to be informative?
- How strong is our confidence in attribution?
Choosing the Right Evaluation Method

Further Reading
Readers who would like to explore evaluation methods in greater depth may find the following resources useful:
- The Magenta Book: Guidance for Evaluation – UK Government guidance on designing, conducting and using evaluations.
- Magenta Book – Annex A: Analytical Methods for Use Within an Evaluation – An overview of the strengths, limitations and applications of common evaluation methods.