Lightweight notes on user-visible improvements. Detailed engineering changelogs and operator runbooks stay internal.
Ask the review assistant why a record was screened out
Screening decisions, with their evidence
The “Ask” assistant can now explain a screening decision. Ask why a record was excluded, included or left unsure — by its title, PubMed id or DOI — and it reports the recorded decision, how confident the screening was, and for each screening question whether the abstract met it, the reason given, and the passage quoted from the abstract. It reports what was recorded; it does not re-read the abstract and reason on its own.
It can also list records by screening status and a title keyword — “which excluded records mention preemptive?” — with the recorded decision and the questions each record did not meet alongside every title, so you can pick one and ask about it.
When asked which screening questions the review uses, the assistant now returns the questions saved for the review’s own framework — Population, Exposure, Study design, Comparator, Outcome and Context for a PEO review — matching the Screening sidebar. Previously it could quote generic placeholder questions on non-PICO reviews and wrongly report a discrepancy.
Ask the configuration assistant about a paper you upload
Attach a review and ask what it did
On the Configure page you can now attach a PDF — a published review you are reproducing, a protocol, or a search appendix — and ask the assistant about it. Attach it from the paperclip in the message box; it takes a minute or two to prepare, and the chip above the box tells you when it is ready.
The assistant does not read the whole paper. It waits for you to ask, then looks up only what you asked for: which outcomes and effect measures the review used, its eligibility criteria, which databases it searched and on what date, which risk-of-bias checklist it applied, or which analysis model it fitted. Tap one of the suggested questions or write your own.
Answers cite the page they came from, and say so plainly when the paper does not report something rather than filling in a plausible value.
Nothing in your configuration changes until you say so. The assistant lists the exact fields and values it proposes and waits for your go-ahead, and anything it applies goes through the same checks as a change you make by hand.
The review assistant is easier to use in a narrow panel
Nothing is hidden at the edge of the panel
Every starting suggestion is now reachable. When more suggestions exist than fit on one line, the row shows a scrollbar instead of quietly cutting the last ones off.
The support button no longer sits on top of the assistant’s send button — it now steps aside while the assistant panel is open, so a question can be sent with the mouse as well as with the Enter key.
A wide table in an answer keeps all of its columns. The table scrolls sideways within the panel with a visible scrollbar, and the expand control still opens it full width.
A wrong number can be corrected against its own source quote
Narrower, safer automatic corrections
When a quality audit finds a value that contradicts the quotation it came from — a misplaced decimal point, or one outcome’s figure recorded against another — it can now correct just that value instead of re-reading the whole study. Only the named figures change; the study groups, units and everything else stay exactly as they were.
A correction of this kind is only allowed to use a number that appears in the study quotation already stored alongside the value, so it can restore what the source says but never introduce a figure of its own. If the number is not there, or the finding does not point to a single row, it comes to you as a suggestion instead.
Corrections still arrive unconfirmed and are re-checked afterwards, and values you have validated or edited are never touched.
Audit results on the Review page read like findings, not notes
Clearer audit summaries
Audit summaries and the reasoning behind each proposed change are now laid out properly — headings, lists and emphasis appear as formatting instead of as stray punctuation in a wall of text. Long passages collapse to a few lines with a Show more control, so a card stays scannable without hiding anything.
Evidence behind a finding now leads with where it was found — a table, a row, a page in the study — described in review terms rather than by the internal name of the data the audit read.
Each proposed or automatically applied change is named by what it does, for example correcting reversed study groups or re-reading an outcome from the paper.
The audit sequence is visible
When an audit applies fixes and then checks its own work, the audit card now shows that sequence: what each round found, how many fixes it applied, how many passed their follow-up check, and how many were handed to you instead.
The proposal list separates what needs your decision from what was already applied, and puts fixes whose follow-up check failed at the top — those are the ones carrying the most context for your judgement.
Discover finds substantially more papers, and its similar-papers mode works again
Citation chasing surfaces more of what you are missing
Discover no longer requires a paper to be connected to two or more of your seeds before it will show it. On a real review, twenty seeds surfaced 86 new papers instead of 15 — and fewer of them were papers already in the review, because demanding two connections mostly re-finds the well-connected studies you already have. How many seeds a paper connects to is still shown, and still used to rank results.
The similar-papers mode returned nothing at all unless you happened to lower that threshold by hand. It now works with the settings it ships with.
Results are marked when the paper is already in your review, so a list of suggestions no longer mixes new findings with ones you have seen without saying which is which.
When a seed paper cannot be found in the citation database, the warning now names the paper and explains that its references and citing papers were left out of the search — previously it showed an internal identifier that meant nothing.
Deep-audit results are easier to find
Audit results surface on your overview
When an automated quality audit of your extracted data finishes, your project overview now points you to its results on the Review page — even when the audit proposed no changes. Previously a clean or observation-only audit finished silently.
The Review page audit card now also lists any fixes the audit applied automatically, with the study affected and the follow-up check each fix passed.
A finished extraction run says what it actually did
A run that ends with studies awaiting your review no longer reads as a failure. The old summary showed "0 success" beside the awaiting-review count even though every study had extracted; the run now says plainly that the studies extracted successfully with some outcomes flagged for review, and the summary shows only the counts that apply.
Re-extraction runs started automatically by a quality audit are now labelled as such, with a note of how many flagged studies they cover. Previously they appeared as an ordinary, unlabelled run over a handful of studies — confusing when you had just extracted many more.
The bulk-extract panel now mentions the parsed full-text requirement only when some studies actually lack one, saying how many are not ready — instead of showing the instruction above every run as if something had gone wrong.
Extraction records the effect measure each study reports, whichever way it differs
Measure disagreements are caught in both directions
Every extraction now records the effect measure the study actually reports, alongside the label as printed. Previously this was only asked for on ecology outcomes, so on clinical ones it was recorded inconsistently or not at all.
A disagreement between the measure a study reports and the measure your outcome is configured for is now flagged whichever way round it falls. Until now only one direction was caught — a hazard ratio on a risk-ratio outcome — and the reverse went unreported.
Extraction no longer removes an estimate that a study clearly labels. A correctly reported protective odds or risk ratio, given in the text without a table of counts, used to be discarded; it is now kept. Estimates with no measure named anywhere can still be removed, and the reason is recorded against the outcome.
Reported hazard ratios are kept when an outcome expects a different effect measure
The measure a study reports is respected
When a paper reports a hazard ratio for an outcome you set up as a risk or odds ratio, the estimate and its confidence interval are now kept. Previously the value was removed and the outcome was left showing its source quotes with no numbers, which read as a failed extraction rather than a mismatch between the paper and the outcome setup.
Any outcome whose reported effect measure disagrees with the one configured for it is now flagged for your review. The estimate is still pooled on the measure you configured, so confirm the measure — or set the outcome up as time-to-event — before relying on the pooled result.
Extraction still removes an estimate that appears to belong to a different outcome, and now records why against that outcome. A blank field can be told apart from one the paper never reported.
Screening distinguishes studies that analyse your comparison from studies that merely mention it
Sharper eligibility screening
Every screening now records whether a study actually analyses the comparison your review asks about — reporting outcomes separately for the compared groups or estimating an effect of the contrast — or only mentions it in passing, for example as a baseline characteristic or a subgroup label. The judgement is visible on each screening decision.
A new protocol option lets comparison-based reviews require the analysed comparison for inclusion. With it enabled, studies where the comparison is only mentioned are excluded, and studies where the text does not reveal which are held for your review. This is designed for questions whose comparison vocabulary appears in many off-topic abstracts — timing-of-treatment and prognostic-factor reviews especially. Ask the configuration assistant to enable it.
Eligibility criteria written as topic labels (for example, a bare description of the comparison group) are now screened as explicit yes/no questions, so a study can fail them. Previously a label could be satisfied by any abstract that used similar vocabulary; the configuration assistant now also flags label-style criteria while you set up the protocol and suggests question phrasing.
Screening handles abstract-less records and checks its own reasoning
Screening decisions you can trust
Records with no abstract in any source — common for older literature — are no longer parked automatically. Screening now reads the title together with the publication-type labels the record is indexed under (for example, "Randomized Controlled Trial"), and marks any decision made this way as title-based so you can confirm it.
A confident exclusion is no longer honoured when its stated reason is contradicted by the screening’s own assessment of the paper — for example, excluding a record as secondary literature while assessing it as a primary study report. Such records are held for your review instead of being silently dropped.
When a record cannot be screened at all, its eligibility criteria are now recorded as unknown rather than as explicitly "not met", so review queues and exports no longer show judgements that were never made.
Screening records now keep both the original suggestion and the decision actually applied, so a suggestion the pipeline declined to act on can no longer be mistaken for the final verdict.
Reference deduplication
Citations without a title no longer merge just because they share a publication year. Title-based duplicate matching now requires an actual title, keeping citation-only records from the same year distinct.
Portfolio conversations stay visible and load reliably
Fixes for asking across your reviews
With many reviews in scope, the conversation is no longer pushed off the screen: the "Reviews in scope" list stays compact and scrollable, and the conversation area now always keeps room for your question and its answer.
The Conversations menu now loads your saved conversations dependably when opened. If the list can’t be fetched — or can’t be refreshed — the menu says so and offers to try again instead of showing an endless loading state.
Higher-recall search strategies, with new checks that prove it
Search queries now match the way papers are actually written
Multi-word concepts in PubMed and Cochrane CENTRAL queries are now matched as near-phrases (the words may appear a few words apart, in any order) instead of only as exact phrases. Exact-phrase matching silently missed studies that phrase the same concept slightly differently; near-phrase matching is strictly more inclusive.
Concepts defined by avoiding, stopping, or scheduling a treatment (for example "…-free" regimens, withdrawal strategies, or preventive versus symptom-triggered timing) are now searched from both sides of the comparison, because papers describe the same trial from either perspective.
When a review’s search window ends many years in the past, the search assistant now includes the terminology of that era, not only today’s consensus terms.
OpenAlex searches now combine synonyms correctly. Previously, adding synonyms could silently narrow an OpenAlex search instead of widening it; the same vocabulary now returns the full breadth it should. OpenAlex results also respect the review’s date range consistently and are no longer restricted to English unless the protocol asks for it.
New safeguards on every search build
Every built strategy is checked term-by-term against the search plan: if vocabulary you specified did not make it into a query, the build reports exactly which terms were dropped instead of discarding them silently.
If you provide known-relevant papers, the built strategy must actually find them on its own. A strategy that cannot retrieve your known papers now fails the build with a clear explanation, rather than being saved with a hidden gap.
Broad search mode now broadens the structure of the queries themselves — wider near-phrase matching and fewer restrictive blocks — not just the size limits.
The notification menu now scrolls through a long backlog
Older unread notifications are reachable again
When you had more unread notifications than fit in the bell menu, the list stopped at the first few and the rest could not be reached. The list now scrolls, so every notification is readable down to the oldest.
Configuration changes are validated before they are saved
Invalid setups are caught at save time, not at screening time
Every change to a review’s configuration — whether made by the setup assistant or through the configuration forms — is now checked for completeness and consistency before it is stored. Common slips, such as an outcome missing its name or an out-of-range outcome classification, are corrected automatically where possible.
When a change cannot be repaired automatically, it is refused with a clear message instead of being stored silently. Previously, an invalid configuration could be saved and only surface later as a screening run that appeared to start but never produced any decisions.
Cross-review conversations are saved — pick up where you left off
Conversations across your reviews now persist
Asking questions across your systematic reviews now keeps the conversation: reloading the page resumes where you left off, and a new Conversations menu lists your recent sessions so you can reopen any of them, delete ones you no longer need, or start a fresh one.
Reopening a conversation also restores which reviews it was about, so the answers keep their original context. Changing the review selection mid-conversation now shows a brief note that earlier answers reflect the previous selection.
An answer keeps working even if you close the tab mid-reply: reopening the conversation shows it still in progress and fills in the finished reply automatically.
After a data lookup, suggested follow-ups (like turning the result into a chart) now appear under the reply on this surface too.
Sensible limits for large portfolios
One conversation can span up to 50 reviews. If you have more, the most recent 50 are selected to start and you can adjust the selection freely.
Very rapid-fire questions are briefly rate-limited, with a clear message asking you to wait a moment and try again.
The assistant on the analysis page keeps its conversation too
The statistics assistant on an analysis’s home page now shares one saved conversation with the statistics workspace: reload the page and the chat is still there, and a question started on the home page can be continued in the workspace (and vice versa).
When a background analysis run finishes, the home-page chat now fills in the reply instead of only announcing that results are ready elsewhere.
If statistical computation is temporarily unavailable, the assistant now says so immediately instead of appearing to work on an answer that never arrives.
A hiccup loading the assistant no longer takes down the whole analysis page — the rest of the page stays available.
Configure appraisal instruments from the analysis settings page, and export risk-of-bias tables to Word
Instrument configuration without editing settings by hand
The analysis configuration page gains a Risk of bias section: choose which appraisal instruments to run, pick variants where an instrument has more than one published form, and fill in instrument pre-specifications — such as the ROBINS-I target trial and confounder list — with inline validation.
Choose how studies are routed to instruments: by detected study design (each study gets the instruments applicable to it), or every configured instrument on every study for concordance comparisons. A read-only coverage hint shows which detected designs your chosen instruments cover, and warns when a design is left without an applicable instrument.
Instrument variants still awaiting verification against their source publication are badged, and a strictness toggle can block batch runs until every configured variant is confirmed.
Risk-of-bias and concordance tables in Word exports
Report downloads in Word format now include a risk-of-bias appendix: one table per configured instrument with per-domain judgements or scores, overall ratings or quality bands, and footnotes marking AI-judged, human-validated entries and imported published assessments.
When a concordance comparison exists, the export also carries the study-by-rater matrix and the pairwise agreement table (paired studies, percent agreement, and Cohen’s kappa with confidence intervals). The Concordance tab gains its own Word download alongside the existing spreadsheet one.
Batch runs are easier to operate
Multi-instrument runs now show progress per instrument in the run monitor, so a stalled or failing instrument is visible at a glance instead of hiding inside one blended bar.
When a run would exceed your usage allowance, it is refused up front with a cost estimate — before any work is queued — instead of after partial setup.
The 2011 Cochrane risk-of-bias tool gains an optional outcome-split form: blinding and incomplete-data entries can be judged separately for each outcome (for example subjective versus objective outcomes), matching reviews that split these entries.
Risk-of-bias assessment now supports multiple appraisal instruments
Four new critical-appraisal instruments
Alongside Cochrane RoB 2.0, analyses can now be configured to appraise studies with the Jadad scale, the original Cochrane risk-of-bias tool (2011), the Downs & Black checklist (including its modified 28-point form), and ROBINS-I for non-randomised studies — so observational studies are no longer skipped when a suitable instrument is configured.
Each instrument keeps its own published scale and scoring: numeric scores with quality bands for Jadad and Downs & Black, low/unclear/high entries for the 2011 Cochrane tool, and the five-level ROBINS-I scale judged against a pre-specified target trial and confounder list.
Where an instrument publishes no scoring algorithm, proposals are clearly marked as AI-judged and require human validation before they can influence certainty ratings.
Risk-of-bias matrix across instruments
When more than one instrument is configured, the risk-of-bias page shows a tab per instrument, with columns and colours drawn from each instrument’s own domains, scales, and score bands.
Published assessments from an original review can be imported from a spreadsheet, side by side with the platform’s own assessments — the foundation for concordance comparisons between tools and between reviewers.
Concordance between appraisal tools
When two or more appraisal perspectives have rated the same studies — different instruments, or the platform’s validated assessments alongside an imported published set — a Concordance tab compares them: a study-by-rater matrix on a shared four-level scale, per-pair percent agreement and Cohen’s kappa with confidence intervals, and a spreadsheet download of both tables.
The extraction workspace reads better — in plain language, in dark mode, and from the keyboard
The provenance panel speaks plain language
The pipeline summary in the evidence panel now describes each extraction run in plain terms — how the values were captured, when, and their review status — instead of internal system identifiers. Individual pipeline steps got the same treatment, and the panel no longer spills past the edge of a narrow sidebar.
Outcomes without a reported timepoint now say "Not specified" instead of presenting a placeholder as if it were a real timepoint.
A sentence that supplies several extracted values — an effect estimate and both confidence bounds, say — now appears once in the quoted-evidence list, annotated with the values it supports, instead of repeating once per value.
Dark mode covers the whole page
Editing fields, the arm-data grid, the run monitor, the extraction summary rail, and the review queue now follow your theme. Previously several of these stayed light on an otherwise dark page.
Keyboard and screen-reader access
The study list can be navigated and activated entirely from the keyboard, form fields are properly labelled, outcome tabs announce which one is selected, and status dots carry text for screen readers rather than colour alone.
Hover explanations throughout the extraction workspace
Controls that could not explain themselves now do. Icon-only buttons say what they collapse or refresh; a greyed-out action says why it is unavailable — most usefully, that automatic extraction needs a parsed full-text document first.
The vocabulary in the evidence panel is explained where you meet it: what each step of the extraction summary counts, what a review verdict means, what a confidence figure describes, and where a value was read from.
Deliberately sparing: anything already carrying a visible explanation, and any control whose label is self-explanatory, has no tooltip.
A tidier workspace with clearer actions
The study title no longer repeats as a large multi-line heading inside the editor — it shows once in the page header, and the editor header carries a compact single line with the year, journal, and PMID instead.
Action labels now say what they act on: "Validate study" (now the highlighted action once a study is reviewed), "Extract all", "Extract filtered" (shown only while a filter narrows the list), and "Save instructions" — which sits with the instructions editor and enables only when there is something to save.
One run summary line reports studies awaiting review instead of three; a finished run says "run finished" rather than "complete" when studies still need review; and an empty editor explains how to pick a study instead of repeating "select a study" in three places.
Extraction says whether each quoted sentence is really in the paper
Every quoted excerpt now carries a verdict, not a disclaimer
Each excerpt in the extraction panel says whether that sentence was found in the document it is attributed to, and on which page. An excerpt that could not be found is marked clearly; one that matched only approximately says so, with how close the match was. Excerpts nothing has examined stay unmarked rather than being presented as confirmed.
Where a document was never stored in a form we can search — some tables, in particular — the panel says there is nothing to check against. That is deliberately worded differently from "not found": it is a gap on our side, and it should not read as a problem with the extraction.
You can settle a disputed excerpt yourself
When an excerpt cannot be located, you can record whether it is in the paper. This is worth doing: a sentence often fails an automatic check for reasons that have nothing to do with whether it is real — italics lost when a PDF is converted to text, a sentence split across a column break, a table flattened into a line. Your answer clears the flag for everyone on the analysis and survives the document being re-processed.
Recording that an excerpt is missing never changes or deletes the extracted number. It records the doubt and leaves the decision to you.
The review assistants present their answers and their work more clearly
Replies render as formatted text
Answers from the qualitative synthesis assistant and the configuration assistant now display with proper formatting — lists, tables, headings, and emphasis — instead of as a single block of plain text. Longer summaries, such as a drafted thematic narrative, are substantially easier to read.
You can see what the assistants did, step by step
Each step the synthesis assistant takes now appears as its own card with a clear running, completed, or failed state, and a run of several steps folds into a one-line summary once it finishes cleanly. Themes, framework matrices, and code lists with content stay visible — completed results are not folded away with the steps that produced them.
The configuration assistant now shows each configuration action as it happens, with a clear success or failure state, and confirms each update it applies.
Every reply can be copied with one click, and the synthesis assistant’s latest reply can be regenerated.
A message that fails to send is no longer lost
If sending fails — a dropped connection, for instance — your message now returns to the box so you can try again, instead of disappearing. Interrupting an answer part-way also no longer disrupts the rest of the conversation.
Results are presented as tidy, self-describing cards
Charts, drafted themes, framework matrices, and coding summaries now sit in a consistent card with a title, a one-line summary, and quick actions — such as copying the underlying data or jumping to the full-size plot.
Supporting quotes in drafted themes cite their study by author and year — with a hover preview showing the full title — and synthesis replies that draw on study evidence list the studies it came from. Framework matrices label their rows the same way.
Suggested follow-up questions in the analysis assistants now sit in a single scrollable row instead of stacking, so they take less space in narrow panels.
Data Lineage now tells you when it cannot back up a quoted value
A quoted excerpt is only presented as a citation when it is really there
When you open a value in Data Lineage, the excerpt shown is now checked against the document it is attributed to. If the text cannot be found there, the excerpt is marked as unconfirmed and carries a short caveat, rather than being presented as a verified quotation. This check runs on values extracted previously as well, so it applies to reviews you have already completed.
This matters most for studies that come with several documents — an article plus its protocol or supplementary material. A value read from one of them could previously be attributed to another, and the panel gave no sign that anything was wrong.
Newly extracted values record where they were read from
Extraction now records which specific piece of evidence each number came from, and Data Lineage opens that document at the relevant page. Where the source cannot be pinned down, the panel says so plainly instead of opening the document at its first page, which could look like the value had been found there.
Custom R analyses finish more often, and explain themselves when they do not
Results are no longer discarded after the analysis has already succeeded
A custom R run has to hand its results back in a particular structure before they can be shown alongside your curated results. The assistant was told which pieces to return but not how each one had to be shaped, so a run whose statistics were entirely correct could still be turned away at the last step — and then turned away a second time on the same misunderstanding. The required structure is now set out in full, with a worked example, before any code is written.
When a run is still turned away, every problem with it is now reported at once instead of one layer at a time. Previously the assistant could correct everything it had been shown, only to be told about the next problem on the following attempt, and it is now given more than one chance to get there.
A failed run tells you what happened
A custom run that could not be accepted used to show a one-line heading with an empty result underneath, even though the specific reasons were available. Those reasons now appear in the run card, so you can see exactly what the assistant is working from — and whether it is something worth intervening on.
Time spent on failed runs is now counted
Workspace time used by a custom run that ends in an error now counts towards your review usage in the same way a successful run does. It previously went unrecorded, which meant a run that kept failing could never reach the spending limit that exists to stop exactly that.
Screening is stricter about what it decides on its own
A rationale that contradicts its own answer now goes to review
Screening records an answer for each eligibility criterion alongside the reasoning behind it. When that reasoning argues the opposite of the answer — marking the population criterion as met, for instance, while explaining that the study organism falls outside the population you asked for — the record is now held as unsure for you to settle, instead of being decided on the answer alone.
This mainly affects reviews whose population is defined by category rather than by degree, such as a specific taxonomic group. Studies of a neighbouring group could previously be admitted as a partial match even where the reasoning said plainly that they did not qualify.
Screening has also been made stricter about the distinction behind that failure: an unconfirmed detail is still a match with a note, but a study whose subject falls outside the categories you named is a straightforward mismatch.
Records with no abstract are no longer screened on a citation
Some records carry only a bibliographic citation — journal, authors, affiliations and identifiers — with no abstract text at all. These are long enough to look like an abstract, and were previously screened as though they were one, producing an eligibility decision drawn from a title and an author list.
Such a record is now recognised as having no abstract. Other sources are tried first, and if none supplies the abstract the record is held for review rather than decided.
A confident exclusion is honoured the same way in every review type
Where screening was confident a record should be excluded for a reason outside your eligibility criteria — the wrong study design, laboratory-only work, an off-target exposure — exposure and outcome style reviews used to park it as unsure precisely because it matched your criteria so well. Other review types already acted on it. That inconsistency is resolved: a confident exclusion is now honoured across all of them, and a low-confidence one is still held for you to decide.
Across existing reviews this affects a small number of records, all of them ones screening had already judged ineligible on design or scope. Anything screening was unsure about stays unsure.
Papers that re-report other studies are excluded automatically
Screening now asks whether a paper reports its own participants or re-reports other published studies. Systematic reviews, meta-analyses, network meta-analyses, pooled analyses, narrative reviews, editorials, practice guidelines, regulatory approval summaries and cost-effectiveness evaluations are excluded on that basis. Previously they could pass whenever the population, intervention and comparator matched your question — which they routinely do, because a review of your exact question does describe your exact population.
This matters most on well-studied topics that already carry a large review literature. There, reviews and meta-analyses can account for a quarter of everything that passes screening, and because several of them summarise the same landmark trial, one trial can reach data extraction under half a dozen separate records — each carrying the same result. Pooling that counts a single trial many times over and tightens the confidence interval around it.
A subgroup analysis or a long-term follow-up of a named trial still counts as a primary report and is kept, because it reports that trial’s own participants rather than someone else’s.
Nothing is excluded on a guess: where a paper does not make its type clear, it is screened on your eligibility criteria exactly as before.
Excluded papers are labelled “Secondary literature” on the Screening page, so you can see what was filtered out and put back anything you want to keep.
This applies to screening from now on. Records you have already screened keep their current decision until you screen them again.
See how well your search matches before you run it
Relevance of a drafted strategy is now shown on the Search page
When a strategy is drafted, a sample of what it finds is checked against your review question. The Search page now shows, per source, what share of that sample actually matched — so you can spot a strategy that is off target before running the full search, while changing the query is still cheap.
The sample is small — around fifteen records per source — so treat it as a smoke test rather than a score. It reliably catches a badly aimed strategy, but a difference of two or three studies moves the figure by more than ten points, so it cannot separate two similar strategies. Act on large differences and on the reason given for the misses.
Where results mostly miss for one reason — the wrong study type, or the wrong population — that reason is called out alongside the figures, so you know which part of the query to tighten.
If the relevance check could not be completed for a source, it now says “not measured” instead of showing a score. A strategy that was never assessed no longer looks like a strategy that scored badly, so you will not narrow a good query on the strength of a check that never ran.
If you edit the query after a check has run, the figures are marked as measured for a previous query rather than presented as current.
Data extraction says so when your plan allowance runs out
A blocked extraction now fails loudly instead of silently
Starting data extraction with your monthly allowance already used up now stops immediately with the plan-limit message and an upgrade prompt. Previously it looked like it had started and then sat at 0% indefinitely.
If the allowance runs out part-way through, the extraction is closed out and shown as stopped instead of staying open. Studies it had already finished keep their extracted data; studies it never reached are simply left ready to extract again.
Either way the analysis is released straight away, so you can start extraction again as soon as your allowance resets or your plan is upgraded — no more waiting for support to clear a run that will never finish.
A separate fix means the last extraction your monthly allowance covers now actually runs. It was previously counted twice and refused.
Stalled extractions are cleared automatically
An extraction that stops making progress for a long stretch is now closed out on its own, instead of blocking every later attempt on that analysis.
Your protocol limits always reach the search
Date ranges, exclusions and required terms are now guaranteed
Limits you set in your search plan — the publication year range, terms to exclude, and any extra required terms — are now applied to every generated query, for PubMed, Europe PMC and Cochrane CENTRAL. Previously a drafted strategy could quietly come back without them, so a review scoped to 2015–2024 might search the whole archive.
The same guarantee holds when a query has to be retried or repaired: the protocol limits are re-applied afterwards, so a repair can widen your search terms but can no longer drop your date range or exclusions.
Qualitative reviews keep their study-design filter through the same retries, so a qualitative search no longer risks coming back as a general one.
Known-relevant papers you pin as seeds are still guaranteed to be retrieved even when they fall outside your date range.
Every strategy is checked against the plan it came from
After a strategy is built, it is automatically reconciled against your search plan — confirming each limit actually made it into the final queries, and that every source you selected produced a query rather than only being described. Mismatches are recorded with the strategy.
A guided path from setup to launching the automated workflow
The workflow launcher walks you through search-strategy setup
When the only thing standing between you and the automated workflow is the search strategy, the launcher now offers “Build search strategy” directly — one click takes you to the Search page and starts drafting the strategy for you, instead of a checklist that leaves you to find the right page yourself.
After you review the strategy and lock it, you are brought straight back to the launch chooser with both automated paths ready to start — no retracing your steps between pages.
A banner on the Search page keeps track of the guided visit, so you always know that locking the strategy is the step that returns you to launch. Ordinary visits to the Search page are unaffected.
The automated workflow can draft its own search strategy (rolling out)
Where enabled, the automated pairwise workflow no longer requires a finished search strategy before launch: it drafts the plan, builds and validates the queries, checks the projected result volume against your study allowance, and then pauses for your approval before anything runs.
You review the drafted strategy on the Search page — with the full editor and previews — and locking it is the approval: the run resumes on its own. Declining ends the run cleanly so you can refine and relaunch. Reviewers are notified in-app and by email when a run is waiting.
An automated reviewer also inspects every drafted strategy — sampling the records it would actually retrieve and judging them against your research question — and its assessment is recorded with each run as this capability rolls out.
More reliable strategy drafting, honest failures, and safer saves
Automatic search-strategy drafting is fixed: it now reliably produces rich, AI-drafted synonym lists instead of quietly falling back to a basic keyword plan.
Starting an automated workflow without an approved search strategy now stops up front with a clear explanation, instead of appearing to finish having found nothing.
Edits made in the structured search-plan editor on the Search page are now saved reliably (previously they could be lost without warning), and run-failure notices use plain language throughout.
Analysis runs catch you up after a refresh
The statistics workspace stays in sync with long-running work
If you refresh or leave the statistics workspace while an analysis run is working, the page now shows a live “run in progress” notice when you return, and brings in the finished conversation and results automatically the moment the run completes.
Runs that stop unexpectedly are labeled clearly — including when a run was interrupted by closing the page mid-analysis — instead of leaving a silent, stale screen. Notices for stalled or interrupted runs can be dismissed.
If the assistant paused to ask for your approval and you closed the page, the workspace reminds you that a decision is waiting when you come back.
The lightweight statistics chat in the analysis sidebar shows the same “run in progress” notice and lets you know when a background run finishes.
Custom R analysis (early access)
Run your own R methods from the statistics chat
You can now ask the statistics assistant to run custom R analyses — including methods and packages beyond the curated toolset — against a frozen, versioned copy of your analysis dataset. Enable it with the “Enable custom R analysis” toggle in the statistics workspace; every run waits for your explicit approval before any code executes.
Results come back as clearly-labeled “User-defined” outputs: they appear alongside curated results and plots with their own badge, carry a full package manifest for reproducibility, and are never used in certainty assessments or generated reports.
Bigger jobs — multi-file analyses, third-party package work, simulation pipelines — can be handed to an analysis workbench that works in the background and reports back when finished. Custom analyses can also be frozen as re-runnable recipes so a promoted result can be reproduced on the exact same data later.
Background workbench tasks now show live progress right in the chat: elapsed time, actions completed in the workspace, and a clear done/paused state — no need to ask for status while a longer job runs.
Configure chat keeps the full conversation
Assistant replies persist in your setup conversation
Replies from the setup assistant on the configure page are now reliably saved with your conversation, so the full exchange — your messages and the assistant’s answers — is still there when you come back or reload the page.
Companion extension: capture tools wait for your approval
In the companion browser extension, capture tools now run strictly after you press Approve — previously a tool could begin working as soon as it was proposed. Declining a tool now shows it as skipped in the conversation instead of leaving it stuck on “Preparing…”.
Unlock a locked protocol to make amendments
Re-open a locked protocol when you need to change it
When a protocol is locked, its configuration and search strategy are frozen and can no longer be edited. You can now unlock the protocol from the configure page to re-open it for amendment — the recourse the “protocol is locked” message points you to. Previously a locked protocol could only move forward, with no way to make a correction.
Unlocking asks for a short reason and is recorded, so every re-opening of a locked protocol is attributable. After unlocking you can edit the configuration and lock it again when you are ready; your earlier locked versions are kept.
Available to the right people
Unlocking is available to team admins and lead reviewers — and the owner of a personal review — the same people who can lock a protocol. Reviewers and viewers cannot unlock.
Ecology and evolution synthesis is more complete and clearly reported
More ecology effect measures, reported on their natural scale
Ecology and evolution reviews now support the response ratio, standardised mean difference, correlation, coefficient of variation ratio, and variability ratio, alongside a generic option for pre-computed effects. These can be computed from reported means, standard deviations, and sample sizes (or correlations), or supplied directly.
Pooled results are now reported on each measure’s natural scale with the correct comparison point — ratios return to the ratio scale and correlations to the correlation scale — so estimates are harder to misread.
Dependent effects and moderators are modelled, not just recorded
Multiple effects from the same study (multiple outcomes, shared controls, or repeated measures) are analysed with a multilevel model that accounts for their dependence instead of treating them as independent.
Recorded moderators such as habitat, taxon, or stratum can now be fitted directly to see whether they explain variation, with an overall test of the moderators; the run falls back to the unadjusted result if the moderator model cannot be fitted.
When the information needed to link correlated effects is not available, they are treated as independent and that assumption is now surfaced on the result so the pooled uncertainty is not over-read.
Clearer scope and reporting standards
Ecology reviews follow PRISMA-EcoEvo and ROSES reporting standards and support systematic maps. Certainty stays narrative rather than clinical, and phylogenetic meta-analysis and moderator interaction terms remain out of scope.
Ecology and evolution synthesis is rolling out on a controlled basis to selected reviews while its analysis is validated further, so availability may vary as it expands.
Risk difference outcomes now pool from event counts
Absolute effects, computed like relative ones
Outcomes configured as a risk difference are now calculated directly from the events and totals each study reports and pooled across studies, matching how risk ratios and odds ratios are handled — including studies with zero events in an arm.
Results are reported on the absolute scale, where a difference of 0 means no effect, with no ratio conversion applied.
Previously skipped outcomes may now pool
Risk-difference outcomes that previously reported too few poolable studies will start pooling on their next run. A living review that includes such an outcome may report a change at its next scheduled update — expected, one-time behaviour as the outcome becomes poolable.
More accurate usage tracking across every review stage
Usage counts now reflect what you actually ran
The usage page counts one statistical analysis run per request instead of counting internal processing steps, so your monthly run allowance goes as far as intended.
Search activity is now shown as individual search requests, and AI activity from screening, extraction, appraisal, and synthesis is attributed to your account more completely — including resumed extractions and document retrieval that previously went uncounted.
Fairer review budget tracking
Review budget tracking now captures activity that was previously missed and applies the correct discount for reused prompt content, so budget warnings reflect real spend more faithfully.
Clearer setup for advanced reviews, with more reviewer checkpoints before results are final
Network meta-analysis setup is easier to review
Guided network meta-analysis now gives reviewers clearer readiness checks before a model is run. Treatment definitions, network connectivity, multi-arm study caveats, and inconsistency warnings are easier to review before interpreting league tables or rankings.
Ecology and evolution reviews have clearer scope
Ecology and evolution guidance now covers PECO, PICOC, and PO questions, ecology-style effect measures, dependent effects, multilevel synthesis, and the current controlled-rollout limitations.
Search, screening, and full text are easier to explain
The public docs now describe search coverage, citation chasing, screening status, full-text upload, and browser-companion capture in user-facing language, with a clearer reminder that search results still need screening before they become included evidence.
Review automation stays reviewer-owned
Advanced automation guidance now emphasises reviewer checkpoints: protocol-level changes pause for approval, generated reports are drafts, and risk-of-bias, certainty, qualitative, NMA, and ecology interpretations remain reviewer-owned.
Approved review changes now carry through to the updated result
One approval pass, one updated result
When an automated review pauses for decisions, approved changes can now be applied within the same run so reviewers see the downstream effect without waiting for a separate pass.
Safer handling for outdated appraisals
Outdated risk-of-bias appraisals can be retired after reviewer approval, with the prior state preserved for audit and disclosure.
Protocol-deviation records
Approved changes that affect the review are recorded so teams can disclose what changed, who approved it, and when.
A dedicated place to review proposed protocol-level changes
Review page for proposals
Protocol-level suggestions now gather in one review surface with plain-language rationales, before-and-after previews, and approve/reject decisions.
Clearer ownership
The app makes it clearer when a human decision is waiting and which reviewers can approve the change.
Automated review can check the analysis and propose safe next steps
A second pair of eyes
The automated reviewer can check evidence completeness, stale appraisals, source-backed values, and outcome consistency, then surface findings for human review.
Human approval for protocol changes
Anything that would change the protocol or evidence set remains a proposal until reviewers approve it.
Qualitative synthesis is easier to review and report
Agreement reporting is clearer
Coding agreement now leads with a statistic that remains easier to interpret when codes are sparse, with familiar agreement measures retained for context.
Cleaner qualitative workspaces
Project cards, study flow, codebook coverage, and re-run behaviour were tightened so qualitative and quantitative reviews are easier to distinguish and maintain.
Qualitative and mixed-methods reviews now run end to end
From question to confidence-rated findings
Qualitative reviews can now move from question setup through screening, full-text appraisal, coding, themes, CERQual confidence ratings, and Summary of Qualitative Findings outputs.
Reviewer-owned interpretation
The system can draft the descriptive groundwork, but themes, analytical findings, appraisal decisions, and confidence ratings remain reviewer-owned.
Living reviews are easier to monitor and act on
Clearer monitoring state
Living reviews now surface cycle status, pending studies, alert settings, and evidence-change prompts more clearly in the places reviewers already check.
Faster recovery
Failed cycles provide clearer recovery actions, while evidence-change alerts link reviewers to the relevant cycle context.
Risk of Bias, GRADE, collaboration, and reporting became more connected
Quality and certainty workflows
Risk-of-bias review, Summary of Findings, and certainty workflows were expanded so reviewer-validated judgements can flow into reports more consistently.
Team review operations
Teams gained clearer roles, assignments, notifications, conflict handling, and sign-off workflows for collaborative reviews.
Search and full-text improvements
Search coverage, citation chasing, reference-manager handoff, full-text retrieval, supplement handling, and manual upload workflows were broadened for everyday review work.