From Small Samples to Better Predictions: Data Augmentation in Muscle Disease Classification
DOI:
https://doi.org/10.54103/2282-0930/32090Keywords:
Data augmentation, machine learning, Sarcopenia, Bioelectrical impedence analysis, Muscle ultrasound, Small sample size, ClassificationAbstract
Introduction
Recent advances in genomics, neuroimaging, and other technological methods for biomedical data acquisition have led to a significant increase in the complexity and volume of information available in medical research. In this context, machine learning and deep learning approaches are increasingly used; however, their application in clinical and biomedical settings is often limited by the presence of many variables and relatively small sample sizes. This condition may lead to unstable estimates, reduced model performance, and an increased risk of overfitting. For this reason, data augmentation techniques may represent a useful strategy to artificially increase dataset size and evaluate whether synthetic data expansion can improve predictive model performance.
Objectives
The aim of this study was to compare three data augmentation techniques (RandomOverSampler with shrinkage, Jittering, and SMOTE) to improve performance of a Machine Learning model in muscle disease classification.
Methods
We used bioelectrical impedance analysis (BIA) parameters and muscle thickness data measured by ultrasound collected in the M-Brain project (PNRR-MCNT1-2023-12378447). The original dataset included 50 subjects (20 diagnosed with sarcopenia according to EWGSOP2 criteria, and 30 without) and 17 variables (nine for BIA and eight for ultrasound). At the beginning, we used the real dataset to identify the most effective classifier among five supervised classification algorithms (Logistic Regression with L2 regularization, K-Nearest Neighbors, Support Vector Machine with a radial basis function kernel, Linear Discriminant Analysis, and Random Forest). For each augmentation technique, we performed a classification using the KNN (the best classifier) previously trained on real data, on the augmented dataset. We retained only the synthetic samples that were correctly classified or had a probability of being a true positive equal to 1.0, and rejected the rest. Model validation was performed using Leave-One-Out Cross-Validation. Performance was assessed using confusion matrix, ROC-AUC, balanced accuracy, recall, precision, and F1-score. This procedure was also performed on subsets that included either ultrasound or BIA variables.
Results
The KNN model without data augmentation achieved a ROC-AUC of 0.640, balanced accuracy of 0.650, recall of 0.600, precision of 0.571, and F1-score of 0.585. The application of filtered data augmentation improved model performance. RandomOverSampler with shrinkage achieved a ROC-AUC of 0.817, balanced accuracy of 0.758, recall of 0.741, precision of 0.690, and F1-score of 0.714. Jittering achieved a ROC-AUC of 0.826, balanced accuracy of 0.770, recall of 0.815, precision of 0.667, and F1-score of 0.733. SMOTE achieved a ROC-AUC of 0.830, balanced accuracy of 0.764, recall of 0.778, precision of 0.677, and F1-score of 0.724. Additional analyses performed on subsets that included either ultrasound or BIA variables confirmed an overall improvement in performance metrics compared with the original data.
Conclusions
All the data augmentation techniques considered improved classification performance compared with the analysis performed without augmentation. However, Jittering showed slightly better overall performance. These findings suggest that data augmentation is a suitable strategy for improving classification performance with this type of data, though it must be handled carefully.
Downloads
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Fabiana Pia Vitiello, Chiara Scarfì, Santina Caliri, Danilo Caudo, Davide Sattin, Antonio Musarò, Nunzio Iraci, Christian Lunetta, Maria Cristina De Cola

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
How to Cite
Published 2026-09-22
Funding data
-
NextGenerationEU
Grant numbers PNRR-MCNT1-2023-12378447


