Skip navigation
Journal article

Cochrane evaluation of (semi-)automated review methods: protocol for an adaptive platform study within reviews

part of AI and SEI and Systematic evidence synthesis

This study protocol presents an adaptive approach for evaluating (semi-)automated methods used in Cochrane systematic reviews. It describes how these methods will be assessed across review tasks including screening, data extraction and risk-of-bias assessment.

Biljana Macura / Published on 26 August 2026

Read the paper  Open access

Citation

Gartlehner, G., Banda, S., Callaghan, M., Chase, J.-A., Dobrescu, A., Eisele-Metzger, A., Flemyng, E., Gardner, S., Griebler, U., Helfer, B., Jemiolo, P., Macura, B., Minx, J. C., Noel-Storr, A., Rajabzadeh Tahmasebi, N., Sharifan, A., Meerpohl, J. J., & Thomas, J. (2026). Cochrane evaluation of (semi-)automated review methods: Protocol for an adaptive platform study within reviews. Journal of Clinical Epidemiology, 198, 112390. https://doi.org/10.1016/j.jclinepi.2026.112390

Using AI technology at work

Key messages

  • Responsible AI can improve evidence synthesis efficiency and reduce human error.

  • AI-enabled review software now supports teams across many synthesis tasks.

  • Evaluating AI tools for evidence synthesis remains methodologically challenging.

  • Rapid AI advances require real-world methods to compare tools for evidence synthesis.

Background and Objectives

Artificial intelligence (AI) has the potential to improve the efficiency of evidence synthesis and reduce human error. However, robust methods for evaluating rapidly evolving AI tools within the practical workflows of evidence synthesis remain underdeveloped. This protocol describes a study design for assessing the effectiveness, efficiency, and usability of AI tools in comparison to traditional human-only workflows in the context of Cochrane systematic reviews.

Methods

Members of the Cochrane Evaluation of (Semi-)Automated Review Methods (CESAR) project developed an adaptive platform study-within-a-review design, modeled after clinical platform trials. This design employs a master protocol to concurrently evaluate multiple AI tools (interventions) against a standard human-only process (control) across 3 key review tasks: title and abstract screening, full-text screening, and data extraction. The adaptive framework allows for the addition or removal of AI tools based on interim performance analyses without necessitating a restart of the study. Performance will be assessed using metrics such as accuracy (sensitivity, specificity, precision), efficiency (time on task), response stability, impact of errors, and usability, in alignment with Responsible use of AI in evidence SynthEsis principles.

Results

The study will generate comparative data about the performance and usability of specific AI tools used in a semiautomated or fully automated manner relative to standard human effort. The protocol provides a flexible framework for the assessment of AI tools in evidence synthesis, addressing the limitations of static, one-time evaluations.

Conclusion

This study protocol presents a novel methodological approach to addressing the challenges of evaluating AI tools for evidence syntheses. By validating entire workflows rather than individual technologies, the findings will establish an evidence base for determining the viability of integrating AI into evidence synthesis workflows. The adaptive design of this study is flexible and can be adopted by other investigators, ensuring that the evaluation framework remains relevant as new tools emerge.

Plain Language Summary

Doctors and researchers rely on systematic reviews, which are thorough summaries of all available research on a health topic, to guide decisions about patient care. However, creating these reviews is a slow and demanding process, often taking more than a year to finish. AI tools could help speed up this work and reduce human errors, but there are currently no reliable ways to test how well these tools perform in real-world settings. This article describes the design of a study that will rigorously test how well AI tools perform when used in actual systematic review workflows, specifically within Cochrane Reviews. The study will compare AI-assisted methods with the traditional approach, where 2 trained researchers independently complete each step. It will look at 3 main tasks: choosing which studies might be relevant based on their titles and abstracts, reading the full-text publication to confirm which studies should be included, and extracting important information from those studies. A key strength of this study is its flexible design. Instead of testing just 1 AI tool at a single point in time, the study allows researchers to add or remove AI tools as new ones become available, similar to how some modern drug trials are run. This approach helps the study keep up with the fast pace of AI development. Researchers will assess the AI tools based on their accuracy, the time they save, how consistent their results are, and how easy they are to use. The ultimate goal of this study is to give the research community strong evidence about when and how AI can be safely and effectively used in systematic reviews to help summarize medical research.

Read the paper

Open access

SEI author

Biljana Macura
Biljana Macura

Senior Research Fellow

SEI Headquarters