Agentic AI for Automated Sensitivity Analysis of Missing Not at Random Data in Patient-Reported Outcome Measures: A Simulation Study

Authors

DOI:

https://doi.org/10.54103/2282-0930/31766

Keywords:

Missing data, Sensitivity analysis, MNAR, PROMS, Agentic AI, Pattern-mixture model

Abstract

INTRODUCTION
Patient-Reported Outcome Measures (PROMs) are central to evaluating surgical effectiveness. In the NHS England PROMs programme, only 27% of hip replacement episodes yielded post-operative responses in 2023/24. When missingness depends on the unobserved outcome (Missing Not At Random, MNAR), standard methods yield biased estimates. Sensitivity analysis via pattern-mixture models and tipping-point approaches is recommended (ICH E9(R1)), yet requires expert judgment for parameter specification and interpretation, making it underutilized. Agentic AI systems, capable of autonomous reasoning and tool use, may automate this workflow while preserving statistical rigor.

OBJECTIVES
To develop and evaluate an agentic AI system that autonomously conducts MNAR sensitivity analysis on PROMs data, comparing its accuracy and efficiency against exhaustive grid-based approaches.

METHODS
We used the NHS England PROMs 2023/24 dataset (Open Government Licence v3), selecting N=11,629 primary hip replacements with complete Oxford Hip Score (OHS) responses (mean health gain 22.6, SD=10.2). A marginal- benefit synthetic scenario (gain ≈ 5.5) was generated to test tipping-point identification. Non-response was simulated via a selection model P(R=1|X,Y) = expit(α₀ + α₁X + δY*), with δ controlling MNAR severity and target missingness of 15%, 25%, and 35%.
The agent was built using the open-source Strands Agents SDK with Amazon Bedrock (Claude Sonnet 4). It was equipped with five statistical tools for baseline analyses, pattern-mixture evaluation, grid-based tipping-point search, bisection refinement, and clinical context retrieval. The agent autonomously decides which δ values to evaluate, when to refine, and when to conclude.
The comparator was exhaustive search over a fixed 10-point δ-grid with m=20 imputations and Rubin’s rules. Performance was assessed via concordance, tipping-point accuracy, number of δ-evaluations, and computation time across 30 Monte Carlo replications (Phase 1 and 2) and 90 agent runs per scenario (Phase 3), with n=2,000 bootstrap samples.

RESULTS
On real NHS data at 25% simulated missingness, bias increased monotonically with MNAR severity (MI-MAR bias −1.24 at δ=−1) and was negligible under MAR (bias < 0.04 at δ=0). In the high-effect scenario (gain=22.6), both the exhaustive grid (360 replications) and the agent (90 runs) confirmed robustness in 100% of cases. In the marginal-benefit scenario (gain=5.5), the exhaustive grid found tipping points only at 35% missingness (tipping point −4.03 at true δ=−3.0, found in 97% of reps; tipping point −4.63 at true δ=−2.0, found in 53%).
The agent identified tipping points in 100% of runs. At 35% missing with true δ=−2.0, the agent found mean δ̂_tip=−4.52 (SD=0.47), concordant with the grid, using 5.6 tool calls in 71 seconds. The agent achieved higher detection than the fixed grid (100% vs 53%) because adaptive exploration extends beyond predefined boundaries. The agent consistently provided clinically contextualized interpretation.

CONCLUSIONS
An agentic AI system can autonomously conduct MNAR sensitivity analysis on real PROMs data, identifying tipping points with accuracy concordant to exhaustive methods and higher detection sensitivity due to adaptive exploration. The approach reduces the expertise barrier for routine sensitivity analysis in PROMs programmes where high non-response makes MNAR a concrete concern. Source code is publicly available.

Downloads

Download data is not yet available.

Downloads

Published

2026-09-22

How to Cite

1.
Agentic AI for Automated Sensitivity Analysis of Missing Not at Random Data in Patient-Reported Outcome Measures: A Simulation Study. ebph [Internet]. 2026 Sep. 22 [cited 2026 Sep. 25]; Available from: https://riviste.unimi.it/index.php/ebph/article/view/31766
Received 2026-05-25
Published 2026-09-22