Missing Data Imputation Strategies for Real-World Clinical Registry Data in Patients With Substance Use Disorder: A Comparative Analysis Preceding the Estimation of Cardiovascular Outcomes

Authors

DOI:

https://doi.org/10.54103/2282-0930/32115

Abstract

Introduction

Real-world clinical registry data from addiction treatment centers are a valuable source for investigating the association between substance use disorder (SUD) and cardiovascular disease (CVD), but are frequently affected by missing data arising from irregular attendance and heterogeneous data collection. Inadequate handling of missing data may introduce substantial bias in the estimation of epidemiological associations, particularly when missingness varies markedly across variables.

Aim

To compare the performance of non-parametric missing data imputation methods in a mixed-type real-world clinical registry and evaluate their impact on cardiovascular risk estimates across variables with increasing levels of missingness.

Methods

A retrospective cohort of 4,895 unique patients admitted to the San Patrignano addiction treatment center (Bologna, Italy) between 2015 and 2024 was analyzed using 40 variables (8 continuous, 32 categorical). The outcome variable was withheld from the imputation model throughout.

Three imputation methods were assessed: missForest, kNN, and guided hot-deck, all tuned by grid search. MissForest was evaluated using built-in error estimates: NRMSE for quantitative and PFC for categorical variables, whereas kNN and guided hot-deck were assessed through method-specific proxy indicators. A binary cardiovascular outcome was defined as the presence of cardiovascular disease-related conditions. Odds ratios (ORs) and 95% confidence intervals (95%CI) were estimated using complete-case analysis and after imputation for variables characterized by different levels of missingness: cocaine use (1%), obesity (12%), total cholesterol (39%) and overdose history (73%).

To quantify the sensitivity of epidemiological estimates to the imputation strategy, two summary indicators were calculated:

ΔOR=(OR_max-OR_min)/OR_cc

and

ΔIC=(IC_maxwidth-IC_minwidth)/IC_widthcc

where  and  denote the odds ratio and confidence interval width obtained from the complete-case analysis. All statistical analyses were done in R (version 4.5.0).

Results

kNN showed a more favorable proxy-based profile than guided hot-deck, with lower continuous and categorical proxy values (0.032 and 0.072 vs. 0.091 and 0.226, respectively), while missForest yielded NRMSE=0.604, indicating a relatively high value, and PFC=0.110. The impact of imputation on epidemiological estimates increased with missingness proportion. For low-missingness variables (cocaine 1%, obesity 12%), OR estimates were stable across all methods (OR range: 4.0–4.1; ΔOR=0.26; ΔIC=0.44). For total cholesterol (39% missing), modest differences emerged across imputation strategies (ΔOR=0.62; ΔIC=0.55), although the epidemiological interpretation remained unchanged. The largest variability was observed for overdose history (73% missing), where OR estimates ranged from 1.8 to 3.3 (ΔOR=0.62; ΔIC=0.55). Nevertheless, overdose was consistently identified as a positive risk factor across all methods.

Conclusions

The practical impact of missing data imputation depended strongly on the amount of information recovered from incomplete variables. For low-to-moderate missingness, complete-case and imputed analyses yielded consistent estimates across algorithms. Differences in point estimates and confidence interval widths emerged only above 70% missingness, without changing the direction or substantive interpretation. Because reconstruction-performance metrics may be structurally heterogeneous across methods, they are best restricted to within-method tuning. In this setting, the operational value of imputation was more appropriately assessed through its impact on downstream epidemiological inference, with ΔOR and ΔIC providing an interpretable framework for comparing imputation strategies in real-world registry studies.

Downloads

Download data is not yet available.

Downloads

Published

2026-09-22

How to Cite

1.
Missing Data Imputation Strategies for Real-World Clinical Registry Data in Patients With Substance Use Disorder: A Comparative Analysis Preceding the Estimation of Cardiovascular Outcomes. ebph [Internet]. 2026 Sep. 22 [cited 2026 Sep. 25]; Available from: https://riviste.unimi.it/index.php/ebph/article/view/32115
Received 2026-06-29
Published 2026-09-22