Methodology · qualitative synthesis

Qualitative and mixed-methods synthesis

Axelium synthesises qualitative evidence — interview, focus-group, and process-evaluation studies — alongside trials. It can draft the descriptive groundwork, but the analytical interpretation and every confidence rating stay with the reviewer. The result is a transparent chain from a verbatim quote, through a code and a theme, to a confidence-rated finding fit for a guideline.

NOTE · Current scope

Qualitative and mixed-methods workflows are available where enabled for the workspace. CASP is the current critical-appraisal path for eligible qualitative studies with parsed full text. JBI-style appraisal is a planned extension rather than the current default.

Overview

Many review questions cannot be answered by effect sizes alone. Whether a complex, behavioural intervention works often depends on acceptability, adherence, and the barriers and facilitators participants describe — questions that are inherently qualitative and frequently mixed-methods: trials and qualitative studies in one review. Axelium adds a qualitative stream that runs in parallel with the quantitative one and converges on a single integrated evidence view.

The workflow moves left to right: define the question and best-fit framework, code passages against the source text, reconcile coding between reviewers, develop themes, then rate the confidence of each finding with GRADE-CERQual and assemble a Summary of Qualitative Findings.

The qualitative synthesis flow: frame the question, code passages with verbatim provenance, reconcile between coders, develop themes, then rate each finding's confidence and assemble the Summary of Qualitative Findings.

Question frameworks

Qualitative questions use frameworks built for experience and process rather than for comparison of interventions:

  • PEO (Population, Exposure, Outcome) — for experience and risk-factor questions where there is no comparator arm.
  • SPIDER (Sample, Phenomenon of Interest, Design, Evaluation, Research type) — purpose-built for scoping qualitative and mixed-methods evidence.

In a mixed-methods review the trial stream keeps its PICO question and the qualitative stream carries its own PEO or SPIDER question, so each body of evidence is framed in the terms that suit it.

Thematic and framework synthesis

Two complementary synthesis methods are supported. The choice is yours and is recorded with the review.

Thematic synthesis

Codes are developed from the data and grouped into descriptive themes that stay close to what the studies report, then carried into analytical themes that go beyond the primary studies to answer the review question. This is the inductive, bottom-up route, well suited to questions where you do not want to impose a structure in advance.

Framework synthesis (TDF / COM-B)

A best-fit framework synthesis starts from an existing structure and codes the evidence into its categories, adding emergent codes where the data do not fit. For behaviour-change questions — common in rehabilitation and physiotherapy guidelines — Axelium provides templates for the Theoretical Domains Framework (TDF) and the COM-B model (Capability, Opportunity, Motivation → Behaviour), so findings map directly onto recognised determinants of behaviour.

NOTE · Pick the method to match the question

Thematic synthesis is the default when no prior structure should be imposed. Reach for framework synthesis (TDF or COM-B) when the review question is explicitly about behaviour change, or when a guideline panel expects findings organised against a recognised model.

The coding workflow

Coding is where the qualitative data is turned into structured, analysable units. It revolves around a shared codebook and passage-level coding that is always anchored to the source text.

  • Codebook — a living list of codes with definitions, hierarchy, and origin. It mixes two kinds of code:
    • Deductive codes — defined up front, typically from the chosen framework (the TDF domains or COM-B components).
    • Emergent codes — created during reading when the data say something the codebook does not yet capture.
  • Passage-level coding — codes are applied to specific passages of text, not whole studies, so a single study can speak to several themes.
  • Verbatim source provenance — every coded passage records the exact quote and where it sits in the source document. Each finding therefore traces back through its themes and codes to the verbatim text that supports it, with nothing paraphrased away.
Provenance chain: every analytical finding traces back through its descriptive theme and code to the verbatim source passage that supports it.

Inter-coder agreement and reconciliation

Credible qualitative synthesis never rests on a single reading. An automated run therefore codes every study twice, with two deliberately different analytical lenses — one descriptive, staying close to what participants actually said, and one interpretive, attending to latent meaning — applied independently over the same passages. Human reviewers who code the same material join the comparison as additional coders. Whenever two or more coders — human or machine — have coded the same material, Axelium quantifies how much they agree beyond chance. The appropriate inter-coder agreement statistic is computed over the coded set:

  • Gwet's AC₁ — the headline figure: a prevalence-robust coefficient that stays interpretable when one category dominates the codings.
  • Cohen's κ — chance-corrected agreement between two coders.
  • Fleiss' κ — the extension to three or more coders.
  • Krippendorff's α — a general agreement coefficient that tolerates missing codings and any number of coders.

Axelium leads with AC₁ because coding is sparse: any single code applies to only a small share of passages, so two coders agree overwhelmingly on not applied. Under that skew the κ coefficients can fall to zero or below even when observed agreement is high — the well-known kappa prevalence paradox, which understates the real level of agreement. AC₁'s chance-correction is robust to that imbalance, so it is the figure reported in the workspace; Cohen's and Fleiss' κ and Krippendorff's α are retained alongside it for reference.

Disagreements are surfaced side by side for reconciliation: coders compare their codings on the same passage, discuss, and settle on an agreed code, which updates the codebook. In an automated run, split passages are first negotiated one by one; a split that negotiation cannot resolve is accepted at the majority position and flagged low confidence in the audit trail, so an unattended run completes without silently discarding the disagreement. When you reconcile interactively, those same unresolved splits are escalated to you instead. The agreement coefficient can be recomputed after reconciliation so you can report both the initial and the post-reconciliation values.

The human-owned analytical leap

Qualitative synthesis depends on interpretive judgement that must remain with the reviewer. Axelium draws a deliberate line between the mechanical groundwork it will help with and the interpretation it will not make on your behalf.

The system may draftThe reviewer owns
Segment passages and propose or apply descriptive codes.The analytical leap — turning descriptive themes into analytical findings that answer the review question.
Cluster codes into candidate descriptive themes, populate framework matrices, and propose CERQual component ratings.Every finalised GRADE-CERQual confidence rating and its rationale.
Compute agreement statistics and draft narrative and rationale text for review.Framework selection and the final sign-off on the synthesis.

WARN · Drafts are proposals, never conclusions

A machine-drafted descriptive theme or rationale is a starting point that a reviewer accepts, edits, or rejects. Analytical themes and confidence ratings are never auto-finalised — they advance only when a reviewer sets them. This is a deliberate guardrail, not a feature gap.

GRADE-CERQual confidence

GRADE-CERQual (Confidence in the Evidence from Reviews of Qualitative research) rates how much confidence to place in each individual review finding. Confidence is assessed against four components. After an automated run, each finding arrives with a proposed rating per component — methodological limitations informed by the CASP appraisals — and the reviewer confirms, edits, and finalises each one; a finalised rating is never overwritten by a later run:

ComponentThe question it answers
Methodological limitationsHow well were the studies contributing to the finding conducted?
CoherenceHow clear and well-grounded is the fit between the data and the finding?
Adequacy of dataHow rich and how plentiful is the supporting data across studies?
RelevanceHow well does the supporting evidence apply to the review question's context?

The four component judgements combine into an overall confidence in the finding of high, moderate, low, or very low. Each rating carries its explanation — machine-drafted while it is a proposal, the reviewer's own once finalised — so the basis for the confidence level is explicit and auditable.

NOTE · CERQual's four components are not GRADE's five domains

CERQual rates confidence in qualitative findings and is distinct from the GRADE certainty assessment used for quantitative outcomes. They share the high / moderate / low / very low vocabulary but assess different things — CERQual weighs methodological limitations, coherence, adequacy, and relevance; the quantitative GRADE assessment weighs risk of bias, inconsistency, indirectness, imprecision, and publication bias.

Summary of Qualitative Findings (SoQF)

Once findings are rated, they assemble into a Summary of Qualitative Findings (SoQF) table — the qualitative counterpart to a quantitative Summary of Findings. Each row states a review finding, its CERQual confidence rating, an explanation of that rating, and the studies that contribute to it. The SoQF is the headline deliverable a guideline panel reads, and it can be carried into your report.

Reporting follows recognised qualitative-synthesis standards: the coding and synthesis trail is structured for ENTREQ (Enhancing Transparency in Reporting the synthesis of Qualitative research) and eMERGe (meta-ethnography reporting guidance) alignment, so the path from search through coding to rated findings is transparent end to end.

Qualitative studies are currently critically appraised with CASP (Critical Appraisal Skills Programme), the qualitative equivalent of a risk-of-bias assessment; that appraisal feeds the methodological-limitations component of CERQual. The pipeline drafts a proposed CASP appraisal for each eligible qualitative study with parsed full text — every answer grounded in a verbatim quote from the paper — and a reviewer confirms or edits each one; a reviewer's appraisal is never overwritten. JBI appraisal is not the current default path.

Mixed-methods integration

The point of running both streams is to bring them together. Axelium juxtaposes the two bodies of evidence in one integrated view: the certainty-rated quantitative effects (each pooled estimate with its GRADE certainty) alongside the confidence-rated qualitative findings (each finding with its CERQual confidence). Reading them side by side shows where the qualitative evidence explains, qualifies, or challenges the quantitative result — for example, why an intervention that pools to a clear benefit is poorly adhered to in practice.

Mixed-methods integration: the trial stream contributes certainty-rated pooled effects and the qualitative stream contributes confidence-rated findings, juxtaposed in one integrated evidence view.

For the quantitative side of a mixed-methods review — pooling, effect measures, heterogeneity, risk of bias, and GRADE certainty — see Statistical analysis.

Current scope

Qualitative synthesis covers thematic and framework (TDF / COM-B) synthesis with passage-level coding, inter-coder agreement, GRADE-CERQual confidence, the Summary of Qualitative Findings, and mixed-methods integration. Two methods are intentionally not yet supported:

  • Meta-ethnography — its interpretive translation of concepts across studies is irreducibly a reviewer's task and is deferred rather than partially automated.
  • Synthesis without meta-analysis (SWiM) — a structured way to combine quantitative studies that cannot be pooled. It is a quantitative technique rather than qualitative synthesis and is not yet available.