Resources · validation

Validation, limitations, and reviewer responsibilities

Axelium is built to make systematic reviews more transparent and auditable. It does not remove the need for protocol discipline, methodological judgement, human review, and final author responsibility.

Validation posture

Outputs should be interpreted as evidence-backed workflow artifacts: search candidates, screening suggestions, extracted values, appraisal drafts, model results, certainty drafts, and report drafts. Each stage is designed to be inspectable before it becomes part of the final review.

www.axelium.ai · Analysis
Data lineage panel
Fig 1Lineage views help reviewers trace a result back to studies, values, and source evidence.

Automation and human review

  • AI screening suggestions can be wrong or incomplete; reviewer decisions should control formal inclusion.
  • Extracted values can be affected by table structure, reporting ambiguity, OCR quality, or conflicting sources.
  • Risk-of-bias and certainty assessments require reviewer validation before publication.
  • Generated reports are drafts and must be checked against the protocol and evidence.

Search, screening, and full text

Search coverage depends on source availability, indexing lag, query quality, and access rights. Optional source coverage and citation chasing improve recall but do not guarantee completeness. Full-text retrieval can be blocked by paywalls, publisher gates, or poor metadata; manual upload and the browser companion help resolve these cases when you have lawful access.

Screening checks its own reasoning for consistency before a record is decided automatically. If a screening rationale argues against its own eligibility answer — for example, marking the population criterion as met while explaining that the study organism falls outside it — the record is held as unsure for you to settle. The same applies when an exclusion’s stated reason is contradicted by the screening’s own assessment of the paper — for example, excluding a record as secondary literature while assessing it as a primary study report — rather than letting a confident-sounding but self-contradictory verdict drop an eligible study.

Records with no abstract are screened on what genuinely exists. A bibliographic citation is never treated as if it were an abstract, but when no abstract exists in any source — common for older literature — screening considers the title together with the publication-type labels the record is indexed under (for example, “Randomized Controlled Trial”), and marks any decision made this way as title-based so you can confirm it. A record without even a usable title is held for review unscreened.

Records excluded for their design rather than for one of your eligibility criteria — secondary research, or work screening was confident falls outside your question for another reason — are excluded rather than left in the unsure queue. Screening still holds back anything it is genuinely uncertain about, so the unsure queue reflects real ambiguity rather than decisions it could not file.

Extraction limitations

Complex tables, unusual outcome definitions, supplement-only data, multi-publication study bundles, and inconsistent registry versus publication values all require reviewer attention. Use source links, confidence indicators, and review queues to verify important values before pooling.

A study may report an effect measure that differs from the one you configured for an outcome — a hazard ratio where the outcome is set up as a risk ratio, or the reverse. Either way the reported estimate is kept and the disagreement is flagged for review, but the outcome is still pooled on the measure you configured. Confirm the measure before relying on such a pooled result.

Statistical limitations

  • Small numbers of studies make heterogeneity, asymmetry, and subgroup checks unstable.
  • Model choice should be justified by the protocol and data, not selected after seeing the preferred result.
  • Effect measures, timepoints, and direction of benefit must be harmonised before pooling.
  • Exploratory analyses should be labelled as exploratory in reports.

Custom (user-defined) analyses

  • Custom R analyses run your own methods in a secure workspace against a frozen copy of the dataset. Their results are labelled “User-defined” and are not validated by Axelium's curated statistical pipeline.
  • User-defined results never feed certainty-of-evidence assessments or generated reports; treat them as exploratory until independently verified.
  • Reproducibility depends on the recorded package manifest; methods from third-party or unpublished packages carry that package's own correctness risks.
  • Every custom run requires your explicit approval and executes with network access disabled; only package installation from public R package repositories is permitted, in a separate step.

Network meta-analysis

Axelium supports clinical network meta-analysis with reviewer-approved treatment definitions, readiness checks, and certainty review. Network setup warnings, multi-arm study caveats, sparse evidence, and inconsistency signals should prompt methodological review before results are reported.

WARN · Living NMA boundary

One-off NMA preparation and synthesis should not be described as recurring living-review NMA automation. Treat recurring NMA refresh as a separate capability unless it is explicitly enabled for your review. Scheduled living-review cycles still import, screen, and extract new evidence for network reviews, but pairwise statistical comparison is paused for them so pooled results never mix different treatment comparisons.

Ecology and evolution meta-analysis

Ecology/evolution support is in controlled rollout and currently centres on multilevel synthesis for dependent effects, with an optional additive meta-regression on recorded moderators and narrative certainty. Results are reported on the effect measure's natural scale with its null anchor.

  • Phylogenetic meta-analysis is not supported.
  • Moderator meta-regression is additive only; interaction terms between moderators are not fitted.
  • A total heterogeneity summary with prediction intervals is reported for multilevel models; a partitioned breakdown across variance components is not.
  • When sampling covariances among dependent effects are not supplied, those effects are treated as independent and the pooled uncertainty may be understated; this assumption is surfaced on the result.
  • Clinical GRADE, CINeMA, and ranking-style clinical interpretation are outside the current public scope.

Qualitative and mixed-methods synthesis

CASP appraisal, coding, theme development, and CERQual confidence ratings are reviewer-owned. Automated drafts can accelerate the groundwork, but analytical findings and confidence rationales must be reviewed, edited, and finalised by the review team.

Review automation

Automated review may flag missing data, stale appraisals, unit issues, or protocol-level suggestions. Protocol-level changes should pause for explicit approval. Applied changes should be disclosed when they affect the reported review. A deep audit of extracted data can also repair narrow, well-evidenced issues automatically. Automatic repair is deliberately bounded: it never overwrites values a reviewer has validated or edited, every repair is re-checked before it counts and listed on the review page, and anything that cannot be confirmed — or that would change the protocol — comes to a reviewer. It is not a substitute for reviewer judgement about what the review should include.

Access, teams, and billing

Team roles, billing status, and connected-app authorization affect who can start work, review evidence, approve changes, and capture full text. Keep team membership and connected apps current for long-running reviews.