Data-Driven Identification of Clinical Phenotypes Associated With Early Infection Risk After Liver Transplantation Using the Open-Source Framework PhenoCluster
DOI:
https://doi.org/10.54103/2282-0930/32123Abstract
Introduction. Early infections are a leading cause of morbidity and mortality after liver transplantation, yet risk stratification typically relies on single markers rather than on integrated donor-recipient-perioperative profiles. The data-driven identification of latent subgroups offers a complementary approach, although its clinical application is often limited by the lack of standardized, reproducible pipelines that are robust against data leakage.
Objectives. To identify, through latent class and latent profile analysis (LCA/LPA), clinically interpretable phenotypes associated with early infection risk in a liver transplantation cohort, while validating PhenoCluster, an open-source Python framework that integrates phenotype discovery and downstream inference into a single configurable pipeline.
Methods. PhenoCluster (https://github.com/EttoreRocchi/phenocluster) implements an end-to-end configurable pipeline. We analyzed 553 liver transplant recipients (2018-2024) described by 30 mixed donor, recipient, and perioperative variables. Phenotypes were identified through LCA/LPA with StepMix, handling missing data via full-information maximum likelihood (FIML) and selecting the number of classes using information criteria (BIC, AIC, ICL). To prevent data leakage, the train/test split (80/20) was performed before preprocessing. The analysis comprised three steps: (i) estimation of the latent class model and selection of the number of phenotypes on the training set; (ii) prediction of phenotypes on the test set, to assess robustness on unobserved patients; (iii) re-estimation on the full cohort for the final estimates. Classification quality was assessed using the average posterior probability (AvePP) and stability via consensus clustering. Associations with outcomes were estimated using logistic regression (OR, 95% CI, Benjamini-Hochberg corrected p-values) and survival models (Cox, Kaplan-Meier).
Results. Three clinically coherent phenotypes emerged, along a gradient of liver disease severity. Phenotype A (n=251; median MELD 12) served as the reference (any infection at 30 days: 27%). Phenotype B (n=179; median MELD 27, Child-Pugh B/C 90%, thrombocytopenia, hypoalbuminemia, and frequent CRE colonization) showed the highest infectious burden: an approximately threefold increase in any infection (54%; OR 3.12; 95% CI 2.01-4.84; HR 2.35; 95% CI 1.75-3.16), with excess bacterial (OR 2.70), viral (OR 3.10), and fungal (OR 4.39) risk. Phenotype C (n=123; median MELD 9, Child-Pugh A 84%) showed a selectively bacterial risk (OR 1.64), without excess viral or fungal risk; 30-day mortality did not differ across phenotypes. Classification quality was high (AvePP 0.921; 74.6% with a posterior probability above 0.90; AvePP 0.929 in the test set), and internal validation showed no overfitting (training vs test log-likelihood -23.46 vs -23.51; consensus clustering stability 0.699).
Conclusions. Latent class and latent profile analysis of integrated donor-recipient-perioperative profiles identifies reproducible and clinically interpretable phenotypes that stratify early infection risk after liver transplantation. The PhenoCluster framework, validated here, provides a standardized, reproducible approach that is domain-agnostic and applicable to heterogeneous clinical cohorts. Open-source code: https://github.com/EttoreRocchi/phenocluster.
Downloads
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Ettore Rocchi, Cecilia Bonazzetti, Martina Casarini, Matteo Rinaldi, Cristiana Laici, Antonio Siniscalchi, Matteo Ravaioli, Matteo Cescon, Maria Cristina Morelli, Gastone Castellani, Pierluigi Viale, Maddalena Giannella

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
How to Cite
Published 2026-09-22


