Use of Large Language Models to Assist Risk of Bias Assessment Supported by an Implementation Document
DOI:
https://doi.org/10.54103/2282-0930/32139Abstract
Introduction
Recent studies showed poor performance of large language models (LLM) for assessing risk of bias (RoB) with the RoB2 tool. This is in line with the low reliability that humans have in assessing RoB. However, the use of an implementation document (ID) prepared by expert reviewers – i.e., a standardised document providing clear instructions on how to answer each signalling question – can increase inter-rater agreement, and may improve the performance of LLM in RoB-related tasks.
Aims
To determine the best framework for using LLMs to support human RoB2 assessments, and to evaluate the value of an ID for LLM-assisted evaluation.
Methods
We will include three outcomes from randomised controlled trials (RCTs) across three different systematic reviews in which RoB2 was manually assessed with the guidance of an ID. The final manual assessments will serve as the reference standard. We will prompt LLMs to assess RoB by providing them with (i) the full RoB2 guidance and (ii) the text and supplementary files of the included RCTs. The model output will consist of quotes from the RCT or generated summaries for each signalling question, along with suggested judgements for each applicable signalling question, each domain and overall. Two human reviewers will assess the quality and usefulness of the LLM outputs in answering signalling questions. We will compare the performance of LLMs with and without additional ID guidance, focusing on information extraction quality, accuracy, and reliability of suggested answers and overall judgements. This study will be conducted using locally run LLMs to ensure data privacy and model stability.
Results
This is an ongoing study. We anticipate developing a framework for LLM-supported RoB assessments through the RoB2 tool. We will also report on the perceived quality and usefulness of using LLMs to support these assessments. We expect the agreement between LLM-based outputs and human reference standard to be non-inferior to assessments performed by two humans using an ID.
Conclusions
This study will inform approaches to use LLMs to support RoB assessment. Future studies should focus on a prospective evaluation of this proposed framework to measure efficiency gains.
Downloads
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Cinzia Del Giovane, Manuel Marques da Cruz, Bernardo Sousa-Pinto, Paweł Jemioło, Rafael José Vieira, Silvia Minozzi

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
How to Cite
Published 2026-09-22


