Likelihood-Free Bayesian Reinforcement Learning for Response-Adaptive Randomization in Clinical Trials
DOI:
https://doi.org/10.54103/2282-0930/32100Abstract
Introduction
Response-adaptive randomization is a class of adaptive clinical trial designs in which treatment allocation probabilities are updated during the study according to accumulating patient outcomes. By assigning future patients with higher probability to the treatment arm that appear more beneficial, these designs aim to improve trial efficiency while addressing ethical concerns related to patient benefit. However, adaptive allocation requires sequential decision-making under uncertainty, especially when the treatment effect and the response-generating mechanism are only partially known. Bayesian Reinforcement Learning (BRL) provides a natural framework for this problem, but standard approaches usually rely on the availability of an explicit likelihood, which may be unavailable or computationally intractable in realistic clinical settings.
Objectives
We propose a novel likelihood-free full BRL approach for response-adaptive randomization in clinical trials. The objective is to update, online, the posterior distribution of both the unknown environment parameters and the resulting optimal allocation policy, thereby quantifying uncertainty not only on treatment effects but also on the adaptive randomization rule itself.
Methods
The proposed method combines Approximate Bayesian Computation with Iterated Batch Importance Sampling (IBIS), from which the name Likelihood-Free IBIS (LF-IBIS) is derived. At each stage of the trial, the parameters that govern the distribution of the unknown response are propagated and reweighted according to their plausibility, which is assessed using simulated pseudo-trial histories rather than an explicit likelihood. For each posterior particle, the corresponding optimal allocation policy is computed, yielding an empirical posterior distribution over response-adaptive policies. We evaluated the method in a two-arm clinical trial simulation, where patient responses are sequentially observed and allocation probabilities are updated as new outcomes become available. A conjugate Bayesian setting was used as a validation scenario, allowing comparison with the exact posterior distribution.
Results
The likelihood-free procedure closely reproduced the sequential posterior learning of the treatment response probabilities obtained in the validation scenario. The resulting posterior distribution of the optimal response-adaptive policy evolved coherently as patient outcomes accumulated, progressively favouring the treatment arm associated with higher expected benefit while retaining policy uncertainty when evidence was still limited. Additional simulations in settings without closed-form posterior updates showed that the method can maintain online Bayesian learning and adaptive policy updating even when standard likelihood-based BRL is not directly applicable.
Conclusions
LF-IBIS provides a flexible Bayesian framework for response-adaptive randomization in clinical trials, that does not require explicit likelihood specification. The flexibility of the method enables the straightforward incorporation of adverse events and individual patient characteristics into the implicit model. By producing posterior distributions over both model parameters and allocation policies, the method supports transparent uncertainty quantification. This is particularly relevant for modern clinical research settings in which complex data-generating mechanisms, patient-centered outcomes and sequential decision-making require statistically principled, interpretable and ethically aware adaptive methods.
Downloads
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Stefano Masini, Cecilia Viscardi, Michela Baccini

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
How to Cite
Published 2026-09-22


