Methodology · risk of bias
Risk-of-bias appraisal
Axelium appraises included studies with established critical-appraisal instruments, each keeping its own published questions, scale, and scoring. This page catalogues the supported instruments, explains how AI-proposed, human-validated, and imported assessments differ, and describes how to read the concordance comparison between raters.
Instrument catalogue
- Cochrane RoB 2.0 (2016). For individually randomised trials, assessed per outcome. Five domains — the randomisation process, deviations from intended interventions, missing outcome data, measurement of the outcome, and selection of the reported result — each judged low risk, some concerns, or high risk via the tool’s published algorithm over its signalling questions, with an overall judgement derived the same way.
- Cochrane risk-of-bias tool (2011, “RoB 1”). The original domain-based Cochrane tool: sequence generation, allocation concealment, blinding, incomplete outcome data, selective reporting, and other bias, each judged low, unclear, or high risk. The tool itself mandates no overall study-level judgement, so none is invented; reviews that published one can carry it through when their assessments are imported. An optional outcome-split form judges the blinding and incomplete-data entries separately for each outcome — the Handbook permits this split, and reviews commonly use it to separate subjective from objective outcomes.
- Jadad scale (1996). A 0–5 quality score for randomised trials built from randomisation (description and appropriateness), blinding (description and appropriateness), and the accounting of withdrawals and dropouts. Scores of 3 or more are conventionally treated as higher quality. The scale has no middle category — which matters when interpreting agreement statistics against three-level tools (see below).
- Downs & Black checklist (1998). A 27-item checklist for randomised and non-randomised studies covering reporting, external validity, internal validity (bias and confounding), and statistical power. The original form has a 32-point maximum; the widely used modified form simplifies the power item to a single point for a 28-point maximum. Totals are grouped into excellent, good, fair, and poor bands.
- ROBINS-I (2016). For non-randomised studies of interventions, judged against a pre-specified target trial and confounder list. Seven domains — confounding, selection of participants, classification of interventions, deviations from intended interventions, missing data, measurement of outcomes, and selection of the reported result — on a five-level scale: low, moderate, serious, critical, and no information.
Choosing instruments
Which instruments appraise your studies is set on the analysis configuration page, and the configuration assistant now proposes a checklist suited to the study designs your question will retrieve — non-randomised tools for an observational corpus, RoB 2.0 for randomised trials, or both for a mixed one — naming the exact form of each so nothing is left to a default. By default each study is matched to the instruments suited to its design — randomised trials to the randomised-trial tools, non-randomised studies to the non-randomised ones — or you can have every configured instrument attempt every study, which is what a concordance comparison between tools needs. Instruments that require pre-specified review decisions, such as ROBINS-I’s target trial and confounder list, are filled in there too, and are checked before an appraisal can start. Where a tool circulates in more than one form, the exact form is selectable; forms whose transcription has not yet been confirmed against the published source are marked provisional, and you can require confirmed forms only so a batch appraisal refuses to start rather than use one. Results are then read against the form each study was actually appraised with: a shortened form is scored out of its own maximum and banded by its own cut-offs, and if an analysis holds two forms of the same tool each gets its own results view rather than sharing one scale.
The Risk of Bias page always shows the current setup before you generate anything: each chosen instrument appears with the number of included studies it would assess, so a mismatch — an instrument that would assess zero studies, say — is visible before an appraisal starts rather than after. If no instrument has been chosen yet, the page says so, and for a corpus of mostly non-randomised studies it suggests a suitable tool. An “Edit instruments” link takes you straight to the selection card on the configuration page.
AI-proposed, human-validated, and imported assessments
Every assessment carries its provenance, and the three kinds are held to different standards:
- AI-proposed. The platform reads each study’s full text and proposes answers to the instrument’s own questions, with supporting quotes. Where a tool publishes a scoring algorithm (RoB 2.0) or an arithmetic rubric (Jadad, Downs & Black), the judgement or score is computed from those answers by the published rules. Where a tool publishes no algorithm — RoB 1 and ROBINS-I judgements are the assessor’s own — the proposal is clearly marked as AI-judged.
- Human-validated. A reviewer has confirmed or corrected the proposal. Proposals require human validation before they can influence certainty ratings, and only validated assessments participate as a rater in concordance comparisons.
- Imported. Assessments transcribed from a published review, attributed to a reviewer label rather than a platform user. They are treated as the published human judgement — already validated — and form their own rater in concordance comparisons.
WARN · AI proposals are proposals
Interpreting concordance between raters
When at least two raters — different instruments, or platform and published assessments of the same instrument — have rated the same studies, the risk-of-bias page shows a concordance comparison: percent exact agreement and Cohen’s kappa (with a 95% confidence interval) for every rater pair, plus a study-by-rater matrix. Read it with these caveats:
- A shared comparison scale. Each instrument’s results are translated onto a common four-level scale (low / intermediate / high / unclear) so unlike tools can be compared at all. These mappings are reporting conveniences, not equivalences — always consult each instrument’s raw results before drawing conclusions.
- Worst-of-outcomes collapse. Instruments assessed per outcome are collapsed to one study-level rating by taking the worst rating across the study’s outcomes. This is conservative by construction: a study rated low risk on most outcomes but high on one counts as high.
- Kappa’s known caveats. Cohen’s kappa corrects percent agreement for chance, but is unstable at small study counts and can be paradoxically low when most studies share one rating (high agreement, low kappa). Report both statistics, not kappa alone.
- 2×2 comparisons. When either tool in a pair has no middle category (Jadad), agreement is effectively over a two-by-two table and the pair is flagged accordingly — its kappa is not directly comparable to kappa over three or more levels.
- One model is not two raters. Platform assessments of different instruments share a single underlying assessment process, so cross-instrument agreement within the platform is not agreement between independent raters. The methodologically cleaner comparison is platform versus published assessments of the same instrument.
Importing published assessments
To compare against a published review’s own appraisal, import its assessments from a spreadsheet. Each row identifies the study (by PMID, DOI, or exact title), names the appraisal tool (and, where relevant, which form of it), and gives the value for one of the tool’s questions or domains — a judgement level for judgement-based tools, or the item’s points for score-based tools. Downloadable per-instrument templates show the expected layout, and totals and overall judgements are recomputed from the per-item values by each tool’s published rules rather than trusted from the transcription.
Imports are attributed to a reviewer label naming the original review. Re-importing updates earlier imports in place, but never overwrites an assessment a reviewer on your team has since worked on — those are preserved and listed. Study identifiers that match nothing in the analysis are reported back rather than silently skipped.