Use of Large Language Models to Assist Risk of Bias Assessment Supported by an Implementation Document

Authors

DOI:

https://doi.org/10.54103/2282-0930/32139

Abstract

Introduction

Recent studies showed poor performance of large language models (LLM) for assessing risk of bias (RoB) with the RoB2 tool. This is in line with the low reliability that humans have in assessing RoB. However, the use of an implementation document (ID) prepared by expert reviewers – i.e., a standardised document providing clear instructions on how to answer each signalling question – can increase inter-rater agreement, and may improve the performance of LLM in RoB-related tasks.

 

Aims

To determine the best framework for using LLMs to support human RoB2 assessments, and to evaluate the value of an ID for LLM-assisted evaluation.

 

Methods

We will include three outcomes from randomised controlled trials (RCTs) across three different systematic reviews in which RoB2 was manually assessed with the guidance of an ID. The final manual assessments will serve as the reference standard. We will prompt LLMs to assess RoB by providing them with (i) the full RoB2 guidance and (ii) the text and supplementary files of the included RCTs. The model output will consist of quotes from the RCT or generated summaries for each signalling question, along with suggested judgements for each applicable signalling question, each domain and overall. Two human reviewers will assess the quality and usefulness of the LLM outputs in answering signalling questions. We will compare the performance of LLMs with and without additional ID guidance, focusing on information extraction quality, accuracy, and reliability of suggested answers and overall judgements. This study will be conducted using locally run LLMs to ensure data privacy and model stability.

 

Results

This is an ongoing study. We anticipate developing a framework for LLM-supported RoB assessments through the RoB2 tool. We will also report on the perceived quality and usefulness of using LLMs to support these assessments. We expect the agreement between LLM-based outputs and human reference standard to be non-inferior to assessments performed by two humans using an ID.

 

Conclusions

This study will inform approaches to use LLMs to support RoB assessment. Future studies should focus on a prospective evaluation of this proposed framework to measure efficiency gains.

Downloads

Download data is not yet available.

Downloads

Published

2026-09-22

How to Cite

1.
Use of Large Language Models to Assist Risk of Bias Assessment Supported by an Implementation Document. ebph [Internet]. 2026 Sep. 22 [cited 2026 Sep. 25]; Available from: https://riviste.unimi.it/index.php/ebph/article/view/32139
Received 2026-06-29
Published 2026-09-22