Bounded Autonomy: What an Automated Review Must Not Be Allowed to Decide
The useful question is not how much of a review a system can run unattended. It is which decisions it may make on its own — and which must be locked before the run begins.
An earlier note argued that when an AI system produces outputs that have to be defensible, the model should propose, a deterministic algorithm should decide, and a human should validate — the three-layer pattern. That pattern governs a single decision. Is this trial at low risk of bias? Is this outcome eligible? Ask who decided, and the answer should name a published procedure rather than a model’s general impression.
A systematic review is not a single decision. It is a sequence of stages — discovery, screening, retrieval, extraction, synthesis, reporting — and inside those stages, hundreds of decisions of that kind. Governing each one individually still leaves the largest question untouched: who decides what happens next?
That question turns out to matter more than the individual judgements, because the answer determines which evidence reaches the analysis at all. A perfectly governed extraction step is worth very little if something upstream was free to decide, unsupervised, which studies it would run on.
Two kinds of autonomy that keep getting bundled
Discussions of agentic systems reach for a dial. How autonomous is it? Level two, level four, human-in-the-loop, human-on-the-loop. The dial is a poor instrument, because it averages two things that need opposite answers.
May the system carry out a step without being asked each time?
Screen four thousand abstracts against the eligibility criteria. Pull the arms and event counts out of a table in a trial report. Fit the model the protocol specified. Laborious, well-specified, and — crucially — checkable after the fact against the source. Withholding this is a large part of why reviews take as long as they do.
May the system decide which step runs, whether one should be done again, and when the work is finished?
These decisions are not laborious. They take a second each. They are also where the serious failures live, because they determine what the analysis is made of and what it is permitted to conclude.
Why the distinction is load-bearing
The two get bundled because from the outside both look like “the system did it on its own”. They should be unbundled, because the defensible answer for the first is generous and the defensible answer for the second is close to nothing. The stages run automatically. The order they run in, the conditions under which one repeats, and the point at which the work stops are fixed in advance — not chosen, mid-run, by a model reacting to what it has just seen.
The locks come first, and they are not a new idea
Before any stage runs, a reviewer commits to six things: the question, the eligibility criteria, the endpoints, the comparators, the effect measure, and the synthesis method.
None of that is a concession invented for AI. It is a protocol. Prospective registration exists because the methodological community worked out, decades ago, that a decision made after seeing the data is not the same kind of decision as one made before. The reason to fix the effect measure in advance was never that a machine might choose badly. It is that anyone choosing after seeing the results is choosing partly on the results.
What automation changes is not the principle but its enforcement. A protocol in a document is a promise, and promises are kept by discipline and memory across a team working over months. A protocol the system holds as a constraint is a boundary. There is no route from a run to a different effect measure, because nothing downstream of the lock is permitted to select one. The commitment stops relying on anybody remembering they made it.
The re-run problem
Here is the failure that should worry you most, and it is not hallucination.
The model is fitted
The pooled estimate crosses the null. Heterogeneity is uncomfortable — an I² in the seventies.
A general instruction applies
Produce a good analysis. The system also has the ability to re-run earlier stages.
Screening is re-run
Under a slightly different reading of an ambiguous inclusion criterion. Two studies drop out.
The run reports success
I² falls to forty. The estimate is now significant.
Every individual action there is one a competent reviewer might have taken for legitimate reasons, and each one, inspected alone, looks defensible. The result is still worthless, because the criterion was chosen to produce the estimate.
This is the garden of forking paths, mechanised. Gelman and Loken’s observation about human researchers was that you do not need conscious dishonesty to arrive at a result that will not replicate — you only need to make reasonable choices with knowledge of the data. The machine version is worse in one specific respect: it is fast, it never tires of trying another branch, and it emits a complete record of every step, which makes the outcome look more rigorous than a hand-run analysis rather than less. The audit trail becomes camouflage.
So the governing rule cannot be “retry when it seems useful”, and it cannot be delegated to a model’s judgement about whether a retry is warranted — that is exactly the decision under suspicion. It has to be a rule about what kind of fault justifies a retry at all.
Mechanical faults justify a re-run. Statistical results never do.
Mechanical fault
A stage failed to do its job — a PDF that would not parse, a request that timed out, a source briefly unreachable. It produced no answer, or a malformed one. Re-running is not selection; it is doing work that did not get done.
Statistical result
Heterogeneity is high. The estimate is imprecise. The funnel plot is asymmetric. The conclusion is null. These are findings, not faults. Re-running in response to any of them is selection wearing repair’s clothing.
The boundary between those two categories has to be drawn in advance and enforced by something other than the component that wants to retry.
Three properties that make the boundary hold
- Attempts are capped, not adaptive. A stage gets a fixed, small number of attempts. The ceiling does not rise because the system judges the problem tractable, and nothing in the loop can grant itself more.
- A worse result is rejected, not kept. If a permitted retry produces an outcome that is worse by a pre-defined measure, the earlier result stands. Retrying cannot be a one-way ratchet toward whatever the last attempt happened to produce.
- Every traversal is recorded and visible. Not just the attempt that succeeded — all of them, including the ones that changed nothing. A retry budget nobody can see is not a budget. A reviewer looking at a finished run should be able to ask how many times each stage ran, and get a number.
Detect, don’t fix
Consider one of the more consequential errors in data extraction: reading the intervention arm as the control arm. The numbers are all real and all correctly transcribed. They are attached to the wrong groups, so the effect estimate is inverted. A benefit becomes a harm. Nothing about the output looks malformed, which is precisely what makes it dangerous — it will pass every check that only asks whether the values are plausible.
A system can be built to notice this. Call it a probe: re-read the source with the arms deliberately reversed, and see whether that reading fits the paper better than the original did. When it does, something is wrong.
The tempting next move is to have the system swap them back. It should not, and the reason generalises well beyond this one case.
What the probe knows
That two readings of the same source disagree, and that the reversed reading fits the paper better than the original did. That is a high-confidence signal.
What it does not know
Which way round is correct. Acting on the disagreement is a second judgement with its own error rate — and if it is wrong, it introduces the exact error it was built to catch, silently, with a record saying the problem was handled.
So the probe raises the study for a human and stops. The asymmetry justifies itself: a flagged study costs a reviewer two minutes of checking. A silently inverted one costs the review its conclusion, and nobody finds out until someone tries to replicate it.
What makes it stop
Stages run until something interrupts them. The categories below are our way of grouping the behaviour for a reader, not a taxonomy the system recites — but each describes a real reason a run halts and hands back to a person rather than proceeding on an assumption.
Missing or conflicting evidence
A source cannot be retrieved, or two sources for the same trial report different numbers for the same endpoint. Neither is a problem automation should resolve by picking one.
Impossible values
A reported figure fails an arithmetic or logical check — a subgroup larger than the population it came from, an event count above the number randomised, a follow-up point the trial never measured.
Method conflict
The evidence that arrived cannot support the analysis the protocol specified: too few studies for the locked model, or an outcome reported on scales that cannot be harmonised without a judgement call.
Budget ceiling
The run reaches its spending or time limit and parks, rather than quietly degrading into cheaper, worse work.
Reviewer gate
A checkpoint the protocol placed deliberately, where progression requires a human decision regardless of how well the preceding stage went.
A guard that cannot evaluate its own condition fails closed
The budget ceiling is worth drawing out, because it illustrates a principle that applies to every limit in the system. If the check cannot determine how much a run has spent so far, it does not proceed on the reasonable-sounding basis that the answer is probably fine. Cost that has not been reconciled counts as spent. A limit that cannot be read is treated as reached.
This is deliberately the pessimistic reading, and it will occasionally stop a run that had room left. That is the correct trade for a guard: the failure it exists to prevent is unbounded work on a review nobody is watching, and a guard that resolves its own uncertainty in favour of proceeding is not a guard.
What stays human, permanently
Three things do not move, and their staying put is not a limitation waiting to be engineered away.
Certainty judgements
Whether the population studied is close enough to the one you care about, and whether an imprecise estimate is imprecise enough to matter for the decision at hand. Both are irreducibly contextual. A structured draft saves real time; a draft is not a judgement.
Interpretation
What the pooled estimate means for practice, for a guideline, or for a funding decision — an argument made by people who can be questioned about it and who carry the consequences of being wrong.
Sign-off
A run does not finish by concluding. It finishes by arriving at a person. There is no terminal state in which the work is complete and no reviewer has accepted it.
Sign-off in particular is a structural property rather than a policy someone could switch off.
Where these guarantees are weaker than they sound
Every claim above is a design commitment, and design commitments vary in how completely they are enforced. Three are worth naming plainly, since the alternative is letting a word like “validated” cover the gaps.
Provenance checking is not span verification. An extracted value is required to carry a reference back to where it came from, and that reference is checked for presence and shape. Confirming that a quoted passage appears verbatim in the source document is a stronger guarantee, and a different one. We do the first thoroughly. We would rather say so than let the two blur.
Retry limits are per stage, not per review. Each stage has a hard ceiling on attempts. That bounds local thrashing well. It is a weaker constraint on a reviewer who restarts a whole review repeatedly with adjusted inputs — which is why the record of every traversal matters, and why it is visible rather than internal. The defence against motivated restarting is that it leaves marks.
None of this improves the underlying literature. Bounded automation makes a review faster, more consistent, and easier to audit. It does not make an evidence base less biased than it is. If the trials were small, selectively reported, or industry-funded without adequate blinding, a well-governed synthesis will tell you that clearly and quickly — and it will still be a synthesis of those trials.
The standing account of what we check and where the boundaries fall lives in validation & limitations, and a worked example of the whole sequence — including where our numbers diverged from the published original and why — is in the reproduction case study.
The question to ask a vendor
Not “how much can it do on its own?” — every answer to that sounds impressive and none of them is checkable. Ask instead: what is it allowed to decide on its own? Then ask the follow-up that actually separates systems: under what circumstances will it run a stage a second time, and who authorised that?
A system that cannot answer the second question precisely is a system where the answer is “whenever a model felt it should”. That is not automation of a systematic review. It is automation of the thing systematic reviews were invented to prevent.
Further reading
- Gelman A. & Loken E. The garden of forking paths. 2013.
- Cochrane, Campbell, JBI, CEE. Joint statement on responsible AI use in evidence synthesis. 2025.
- Cemri M. et al. Why Do Multi-Agent LLM Systems Fail? (MAST taxonomy). 2025.
- Page M. J. et al. The PRISMA 2020 statement. BMJ 2021.
- European Union. AI Act, Article 12: record-keeping. 2024.
Reading in order? Start with the three-layer pattern, which sets out how a single decision is governed, then return here for how a whole run is. Or open the worked demo and trace a finished review yourself.