Resources · best practices

Best practices

Actionable tips for getting the most out of Axelium at every stage of your systematic review.

Setting up your review

  • Define PICO/PEO before adding studies. The tighter your criteria, the faster screening runs and the lower your unsure rate.
  • Name outcomes precisely. Use “Overall Survival” instead of “OS”. Clear names reduce extraction ambiguity and help Axelium select the appropriate analysis path.
  • Pick the right framework first. Switching from PICO to PEO mid-review forces re-screening. Commit up front.

Search strategy

  • Let Axelium draft, then refine. The first query is a starting point. Use Refine queries to rework the search terms while keeping the rest of your strategy — the year range, exclusions and any terms you added by hand — exactly as you set them.
  • Rebuilding from scratch discards your edits. Rebuild plan + queries from scratch and Reseed plan heuristically both start a brand-new strategy from your research question. Anything you typed in yourself is not recoverable from the question, so it is lost — exclusions most often. Both ask you to confirm first. Once you are happy with a strategy, locking it disables these two actions, which is the simplest way to protect a strategy you have tuned by hand.
  • Edit the plan directly for anything the drafter cannot infer. Edit structured plan opens each part of the strategy — population, exposure or intervention, comparator, outcomes, year range, extra terms and exclusions — as editable fields. Exclusions are matched against titles and abstracts, so prefer distinctive phrases; a common word will also remove relevant studies that happen to mention it. Saving the plan clears the generated queries, so rebuild them before running the search.
  • Treat the relevance check as a smoke test, not a score. Before you run the full search, a small sample from each source — around fifteen records — is screened against your review question, and the share that matched is shown on the Search page. It is good at catching a strategy that is badly off target. It is not precise enough to separate two similar strategies: on a sample that size, a difference of two or three studies moves the figure by more than ten points. Act on large differences and on the reason given for the misses; ignore small ones rather than rebuilding repeatedly to chase them.
  • Screening is where the real filtering happens. A strategy that retrieves some irrelevant records is normal and usually preferable to one that misses eligible studies. Every retrieved record is screened against your full criteria later, so favour recall here and let screening remove what does not belong.

Screening efficiency

  • Phrase eligibility criteria as questions a study can fail. “Does the study itself compare X against Y?” screens far more precisely than a topic label like “Y group”, which any abstract using similar vocabulary can satisfy. Label-style criteria are flagged during protocol setup and screened as questions, but criteria you phrase yourself are always sharper.
  • Require the analysed comparison when its vocabulary is everywhere. Timing-of-treatment and prognostic-factor questions often match abstracts that only mention the comparison as a baseline characteristic or subgroup. For those reviews, enable the option that requires included studies to analyse the comparison — mention-only studies are excluded and ambiguous ones go to review.
  • Write custom screening instructions for your domain. For example, “exclude phase I dose-finding” or “include only RCTs with ≥50 patients”. Domain-specific guidance dramatically cuts the unsure rate.
  • Resolve unsure studies in batches, not one-by-one. Look for patterns — if many unsure studies share the same exclusion reason, refine your PICO or custom instructions instead of deciding each study individually.
  • Aim for <5% unsure before moving to extraction. A large unsure bucket means the meta-analysis is incomplete rather than just underpowered.

Full text and supplements

  • Use Bulk Upload early. Drag all your PDFs in one go — the matcher handles PMID and title-based linking automatically.
  • Don't skip supplements. Many effect sizes — especially subgroup and secondary endpoints — live in supplementary tables. The system parses DOCX and XLSX automatically.
  • Check the “No PDF” filter before extraction. Studies without usable full text are not ready for batch extraction; resolve them before treating extraction coverage as complete.

Extraction quality

  • Run Auto-Extract All Outcomes, then triage. Batch extraction with automated QC is faster than extracting studies one at a time.
  • Trust the confidence badges. Green values rarely need review. Focus your time on amber and red extractions in the Review Queue.
  • Use rerun instructions, not manual edits, for systematic errors. If the extractor keeps pulling from the wrong table, provide guidance like “use Table 2, not Table S3” via rerun instructions and let it re-extract. Manual edits don't improve future runs.
  • Review arm swap flags immediately. A swapped arm inverts the effect direction and will silently corrupt your pooled estimate.

Risk of bias

  • Don't skip the dual-review queue. Risk-of-bias proposals can speed up appraisal, but a second reviewer should confirm or override every domain judgement before you rely on it. Unvalidated judgements should not drive certainty conclusions.
  • Read the evidence spans, not just the traffic light. Each domain judgement links to the quoted PDF text behind it. A quick scan of those spans catches the occasional misread before it shapes your certainty rating.
  • Pick the Domain 2 variant deliberately. Choose “effect of assignment” (intention-to-treat) or “effect of adherence” (per-protocol) per analysis — they answer different questions and can yield different judgements.

Statistical analysis

  • Start with the quick-prompt buttons. They keep common analyses close to the dataset summary and help you avoid running analyses before enough studies are available.
  • Always check I² before interpreting the pooled estimate. High heterogeneity (>75%) means the summary number may be misleading. Run subgroup analysis to investigate sources of variation.
  • Run sensitivity analysis before finalising. Leave-one-out reveals whether a single study drives the result. Estimator comparison confirms robustness to the choice of τ² method.
  • Don't skip publication bias for k ≥ 10. Below 10 studies the tests lack power, but above 10 you should always report it.
  • Spot-check results with Trace Lineage. Before finalising, open the Trace Lineage panel on two or three pooled estimates and follow each back to the source PDF snippet. It's the fastest way to catch an arm swap, a unit mismatch, or a subgroup/overall mix-up.
  • Reach for custom analysis only when the built-in tools can't express the method. The curated tools are deterministic and feed certainty ratings and reports; custom (user-defined) results do not. Use a custom run — or hand a larger job to the Analysis Workbench — for a bespoke model, an uncurated package, or your lab's own method, and treat its output as a supplementary analysis you review, not a replacement for the curated result.
  • Freeze a custom analysis you intend to cite. Promoting it to a recipe captures the exact code, package manifest, and frozen data so the estimate can be reproduced later — the same reproducibility bar you'd expect of a curated run.
  • Pin evidence as you go. It's easier to build the report from pinned artifacts than to scroll back through chat history.

Reporting

  • Pin before you generate. The report generator uses the Evidence Board. Unpinned plots and results won't appear in the final output.
  • Choose target audience up front. Academic, Clinical, Regulatory, and Patient reports differ in tone and detail level. Picking the right audience shapes the entire narrative.
  • Review the PRISMA flow for completeness. Check that unsure studies are resolved to zero and the included count matches your extraction count before generating.
  • Finalise GRADE before exporting the Summary of Findings. Auto-derived certainty is a starting point — review each downgrade factor, record your rationale, and finalise the assessment so the exported SoF table reflects reviewer judgement, not just the defaults.