Machine Learning-Based Classification of Parkinson’s Disease Using Clinical and Biomarker Data: A Leakage-Safe Benchmarking Study With Convergent Feature Ranking

Authors

DOI:

https://doi.org/10.54103/2282-0930/32165

Abstract

Background

Parkinson’s Disease (PD) is a progressive neurodegenerative disorder characterized by nigrostriatal dopaminergic neuronal degeneration and α-synuclein aggregation. Motor symptoms emerge after substantial neurodegeneration, making early diagnosis challenging. A prodromal phase, characterized by non-motor manifestations such as REM sleep behaviour disorder and hyposmia, may precede clinical diagnosis by several years. Cerebrospinal fluid biomarkers and dopamine transporter imaging have shown potential for early risk identification, supporting the application of machine learning (ML) approaches to PD diagnosis and risk stratification.

Objectives

This study aimed to provide a transparent, reproducible and leakage-safe benchmark of supervised ML models for PD classification using clinical and biomarker data, and to identify the most informative features for discriminating PD patients from healthy controls (HC) through convergent feature ranking across multiple selection approaches.

Methods

Data from 58 participants (40 PD, 18 HC) were analysed by integrating demographic, clinical, and biomarker variables, including dichotomous biomarker-status indicators, core neurodegeneration markers (bNfL, bGFAP, bpTau217, bN42, bN33, NSE), vascular/endothelial/hemostasis markers (U-PAR, VEGFA, ICAM-2, VCAM-1, VWF, T-PA, P-selectin, PECAM-1, E-selectin, SELPLG), immune/inflammatory markers and spectroscopy-derived metabolites (NAA, Cho, Glu, Ins) measured in the substantia nigra and globus pallidus bilaterally. Leakage-safe preprocessing comprised feature selection and standardization, with balanced class weighting applied where supported. Eight supervised ML classifiers were evaluated: Logistic Regression, Linear and Radial Basis Function Support Vector Machines (SVM), Random Forest, Gradient Boosting, k-Nearest Neighbours, Decision Tree, and Multi-Layer Perceptron. Hyperparameters were optimized through grid-search with stratified ten-fold cross-validation, using balanced accuracy as the primary optimization criterion and receiver operating characteristic area under the curve (ROC-AUC) as a secondary metric. Analyses were repeated across eight biomarker subsets. Model interpretability was investigated using SHapley Additive exPlanations (SHAP), while feature relevance was quantified through convergent ranking across ANOVA F-test, Mutual Information, Random Forest importance, LASSO coefficients, and principal component analysis (PCA) weighted loadings.

Results

Linear SVM achieved the highest cross-validation balanced accuracy (0.86 ± 0.19) and was selected as the best-performing model. On the independent test set, it achieved a ROC-AUC of 0.88 and a balanced accuracy of 0.79, with F1-scores of 0.71 for HC and 0.82 for PD. SHAP analysis identified MoCA as the most influential predictor. Convergent feature ranking identified IOIT as the dominant predictor in the full-feature set, while MoCA ranked first after excluding dichotomous biomarkers, followed by MMSE and bNfL. Within biomarker domains, bNfL emerged as the leading neurodegeneration biomarker; U-PAR, VEGFA, and ICAM-2 as the most relevant vascular markers, IL2-RA, TNF-R1, IL17RA, and TNFSF13B as the leading immune markers and GP-DX-Glu and SN-DX-Ins as the top spectroscopy derived features. PCA loadings diverged from supervised rankings, indicating limited overlap with discriminative relevance.

Conclusions

Leakage-safe ML pipelines were promising for distinguishing PD patients from HC using multimodal clinical and biomarker data. Although limited by the small sample size, this transparent and reproducible framework supports the use of parsimonious feature sets, with MoCA emerging as a robust predictor. SHAP-based interpretability further enhanced model transparency and the identification of clinically meaningful markers.

Downloads

Download data is not yet available.

Downloads

Published

2026-09-22

How to Cite

1.
Machine Learning-Based Classification of Parkinson’s Disease Using Clinical and Biomarker Data: A Leakage-Safe Benchmarking Study With Convergent Feature Ranking. ebph [Internet]. 2026 Sep. 22 [cited 2026 Sep. 25]; Available from: https://riviste.unimi.it/index.php/ebph/article/view/32165
Received 2026-06-29
Published 2026-09-22

Funding data