Fixed-effect or random-effects: which should I use?
Choose random-effects if you believe the true effect genuinely differs across studies — different populations, doses, or settings — and fixed-effect only if you believe every study estimates one identical true effect. Base the choice on your question and your studies, decided before you see the results, not on the heterogeneity statistic you get back.
What the two models actually assume
The models differ in what they think is going on in the world, not in how conservative they are. That is the part most summaries skip.
- Fixed-effect
- There is one true effect. Every study is measuring that same number, and the only reason the estimates differ is sampling error. Studies are weighted purely by precision, so a large trial can dominate the pooled result almost entirely.
- Random-effects
- True effects vary across studies, and the studies you have are a sample from a distribution of them. You estimate the mean of that distribution and its spread. Weighting is more even, so smaller studies count for relatively more.
The two answer different questions. Fixed-effect answers "what is the effect in these studies?" Random-effects answers "what is the average effect across the range of settings these studies represent?" For most clinical reviews the second is the question you actually have.
Why "I² was high, so I used random-effects" is the wrong reason
It is the most common justification given, and it inverts the logic. Picking a model after seeing the heterogeneity statistic makes the model a result rather than a decision, and it means the same set of studies would have been analysed two different ways depending on which trials happened to be published.
The decision is about whether the studies are conceptually identical replicates. If they enrolled different populations, used different doses, or ran in different care settings, the true effects almost certainly differ — and that remains true whether the heterogeneity statistic comes back at 0% or 80%. Pre-specify the model in the protocol and record the reasoning.
In practice, random-effects is the default for most clinical systematic reviews, because studies of the same intervention are rarely exact replicates of one another.
What changes when you pick random-effects
The confidence interval widens, because it now carries the uncertainty in the between-study variance as well as the sampling error. That is not the model being cautious for its own sake — it is the interval telling you the truth about a set of studies that disagree.
You should also report a prediction interval. The confidence interval describes the average effect; the prediction interval describes the range a new study might plausibly land in. When studies genuinely differ, the prediction interval is usually far wider than readers expect, and it is the honest summary of what you know.
Estimating the between-study variance requires an estimator. Restricted maximum likelihood (REML) is the usual default for continuous and ratio measures and is a defensible choice to pre-specify. With few studies, a Knapp–Hartung adjustment to the confidence interval is worth considering, since it accounts for the uncertainty in that variance estimate.
When should I not pool at all?
Neither model answers this, and it is the more important question. If the studies address meaningfully different questions — different comparators, incompatible outcome definitions, populations that would never appear in the same guideline recommendation — then a pooled number is a precise summary of something nobody asked about.
A structured narrative synthesis, or separate pooled estimates per subgroup with the subgroups pre-specified, is the better answer more often than the literature implies.
Where this answer stops
With fewer than about five studies, the between-study variance is estimated poorly and random-effects intervals can be unreliable in either direction. Neither model rescues a set of studies that should not have been pooled at all.