Skip to content
All guides

How do I pool outcomes measured on different scales?

Effect measures · Updated July 2026

Short answer

Use a standardised mean difference, which rescales each study's effect by its own standard deviation so results from different instruments become comparable. Prefer the small-sample-corrected form (Hedges' g). Only do this when the instruments genuinely measure the same underlying construct — the method cannot tell whether they do.

When you need it

The classic case is a construct with several established instruments — depression severity, pain, quality of life, function — where trials each picked a different one. The raw mean differences are not comparable because the scales have different ranges and different meanings per point.

If every trial used the same instrument, do not standardise. Pool the plain mean difference: it stays in units your reader already understands, and it avoids the sample-spread problem entirely.

Which version to use

Hedges' g
The standardised mean difference with a correction for bias in small samples. This is the sensible default; the correction is negligible in large trials and matters in small ones.
Cohen's d
The uncorrected form. It is biased upward in small samples, which is exactly where meta-analyses of psychological and rehabilitation interventions tend to live.

The 0.2 / 0.5 / 0.8 "small, medium, large" labels were offered as a last resort when nothing better was available. They are not clinical thresholds, and a standardised effect of 0.3 can be decisive or irrelevant depending entirely on the outcome.

Getting the direction and the scale right

Instruments do not agree on which way is better. On some scales a higher score is a worse symptom; on others it is better function. Before pooling, fix a single direction for the construct and flip the sign of any study measured the other way. This is the most common and most silent error in standardised pooling — mixed directions pull the estimate toward zero and look like a null result.

You also need a standard deviation per arm. Trials frequently report a standard error, a confidence interval, or a median with an interquartile range instead. The first two convert exactly; the third requires an estimator whose assumptions about the distribution should be stated, and whose use belongs in your methods section rather than a footnote.

Making it interpretable again

A standardised effect is hard to act on. The usual remedy is to translate the pooled estimate back onto the most familiar instrument by multiplying by a representative standard deviation, and then to compare it against a minimally important difference for that instrument if one exists. Report the translation and the standard deviation you used, so a reader can disagree with your choice rather than guess at it.

Where this answer stops

A standardised mean difference is divided by the standard deviation of the sample it came from, so a trial in a narrowly selected population will report a larger standardised effect than an identical trial in a broad one. Differences in the spread of the enrolled population can masquerade as differences in effectiveness.