Skip to content
All resources
Evidence note6 July 2026·8 min read

When the indirect comparison is the weak link

Regulatory approval lets a medicine be sold. A separate value gate decides whether a health system will pay for it — and that gate increasingly turns on an indirect comparison the trial never ran. Reading across published HTA appraisals, that comparison is, again and again, the step a committee finds fragile.

2 gates

Approval ≠ reimbursement

A medicine can clear the regulator and still be restricted, delayed, or price-capped at the payer gate.

~2 years

Approval → decision

A typical wait from approval to a reimbursement decision — during which new trials keep landing and the evidence base moves.

8 → 102

Single-arm submissions

HTA submissions built on single-arm trials rose roughly ten-fold across the 2010s — each one forcing an indirect comparison at the value gate.

Two gates, two questions

A regulator asks a narrow question: does the medicine work, and is it safe — often measured against placebo. A payer or health technology assessment (HTA) body asks a harder one: is it better than the treatment patients already get, and is it worth the price. Those are different questions, and the second one needs a comparison the first rarely provides.

More and more, the head-to-head trial a payer would want was never run. So the comparison has to be manufactured after the fact — an indirect comparison built from separate studies: a network meta-analysis (NMA), a matching-adjusted indirect comparison (MAIC), a simulated treatment comparison (STC), or, in the worst case, a naïve unadjusted one. That step is the single most contested part of a modern value dossier, and it is largely built in free, open-source statistical code with no productised, inspectable workflow around it.

Why this keeps happening

The volume of hard cases is rising. Submissions built on single-arm trials — where there is no internal control, so a comparison must be constructed indirectly — grew from a handful a year at the start of the 2010s to around a hundred by the end of the decade. At the same time, the methods themselves are a moving target: population-adjusted approaches have consolidated around the NICE Decision Support Unit guidance (TSD 18), while newer extensions such as multilevel network meta-regression keep raising the bar a dossier is measured against. The standard is getting harder at the same moment more launches are forced onto indirect evidence.

What the appraisal record shows

We read a set of published appraisals in which a medicine was recommended only for a narrower population than its licence — “restricted-population” decisions — and looked at what the committee actually said about the evidence. Two things stood out.

First, the indirect comparison was named as a source of the committee’s uncertainty in the clear majority, and in a large share the synthesis itself — not just the underlying data — was the stated locus of the problem. Second, when a specific method was named, the mix was consistent:

Stacked methodsThe common case

Most decisions we read cited more than one indirect-comparison method used together — for example a network meta-analysis alongside a matching-adjusted comparison.

Network meta-analysisMost common single method

Where a committee named one method, a network meta-analysis (NMA) was the most frequent.

Unanchored MAICThe hardest cases

Matching-adjusted indirect comparisons without a common comparator recurred in the appraisals that drew the sharpest critique.

Naïve comparisonRare, sharply criticised

Unadjusted comparisons were uncommon, but drew the strongest objections when a dossier leaned on one.

What committees actually object to

The objections were not really about statistics for their own sake. They clustered into a few recognisable failure modes — each one a place where a comparison quietly assumes something it cannot show:

Unanchored comparison

No common comparator links the trials, so the result rests on an assumption that the treatment effect does not depend on patient characteristics — one committees rarely accept at face value.

Population mismatch

The trial population differs from the decision population. Matching adjusts for measured differences, but not for the ones nobody recorded.

Unexplained heterogeneity

Studies in the network disagree more than chance allows and the source is not explained, so the pooled estimate is hard to trust.

Uncertainty into the model

A fragile comparison feeds an uncertain relative effect into the economic model, and the cost-effectiveness estimate inherits every bit of that uncertainty.

Why a fragile comparison is expensive

When a committee finds the indirect comparison fragile, the verdict is usually not a flat no. It is a restriction to a narrower population, a managed-access arrangement, a price cap, or a delay while more evidence is sought. The uncertainty propagates: an unanchored comparison feeds an uncertain relative effect into the economic model, the cost-effectiveness estimate widens, and the committee hedges. The comparison is a high-leverage step — get it wrong and you lose population, price, or time, not just a p-value.

What evidence teams can do about it

None of these failure modes is a surprise on the day of the committee. Each one can be anticipated, and mostly answered, while the comparison is still being built. Three moves matter most:

1

Build to the standard it is judged against

Pairwise and network meta-analysis, population-adjusted comparisons (MAIC, STC, ML-NMR), and GRADE certainty — with every pooled estimate traceable back to the study it came from.

2

Stress-test before the committee does

The statistics are deterministic and inspectable: open any estimate, change the comparator, the population, or the model, and re-run to see how much the result actually moves.

3

Keep it current

New trials land during the long wait to a decision. A living review re-runs on a schedule and flags when the pooled estimate changes, so a finished comparison does not quietly go stale.

This is what Axelium is built for: the same indirect-comparison methods a committee expects, run on a deterministic engine you can open and interrogate, and kept live as the evidence changes. The point is not to predict a verdict — it is to see the weak point in your own comparison before an assessment group does, and to fix it while there is still time.

How to read these findings

  • This is a non-exhaustive read of public appraisal records — a partial slice, and an absence in it is not evidence that a problem did not exist.
  • We describe patterns in aggregate and name no company. A restricted recommendation reflects a committee’s view of an evidence package on a given day, not a verdict on a sponsor.
  • Where we group what committees said, those groupings are our analytical reading of public documents, not the committees’ own labels.
  • Figures are point-in-time and will shift as more of the public record is read.

Build the comparison that holds up

See how Axelium runs pairwise and network meta-analysis and population-adjusted comparisons on a deterministic, inspectable engine — and keeps a finished comparison current as new evidence lands.

Sources & method: Patterns are our own aggregate reading of public HTA appraisal records (for example, published NICE committee papers) and public methods guidance, including the NICE Decision Support Unit Technical Support Document on population-adjusted indirect comparisons (TSD 18). Named methods (NMA, MAIC, STC, ML-NMR, GRADE) are standard evidence- synthesis terminology. Figures are indicative and point-in-time.

Note: Axelium builds evidence-synthesis software. This is an educational note describing public records in aggregate; it names no company and is not a prediction of any committee’s decision.