Sample Size Determination in Studies Involving Machine Learning Algorithms: A Case Study on Post-Stroke Rehabilitation
DOI:
https://doi.org/10.54103/2282-0930/32117Abstract
Introduction
Sample size determination remains an important methodological challenge in the development of machine learning (ML) prediction models in clinical studies. Unlike classical statistical approaches such as regression-based methods, there is still no widely accepted procedure for estimating the amount of data required to develop reliable ML models, and current practice remains highly variable and largely based on informal or heterogeneous criteria.
Aims
This study had two main aims: 1) to identify and critically compare the methods proposed for sample size determination in ML-based clinical prediction studies through a systematic review; and 2) to apply and compare the methods identified in the systematic review within a real-world case study based on the STRATEGY study, a multicentre Italian cohort originally including 743 post-stroke patients admitted to intensive inpatient rehabilitation.
Methods
The systematic review was conducted using searches in three bibliographic databases: MEDLINE, the Cochrane Library, and EMBASE. In addition, a forward citation search was performed in Scopus to identify studies that applied the identified methodological approaches to justify sample size in their own ML-based analyses.
The identified methods were then applied to the STRATEGY dataset. Model development and evaluation were conducted using a nested validation framework. First, the dataset was split into a development set (80%) and an independent test set (20%). Within the development set, an inner 5-fold cross-validation procedure was used for models training (i.e. SVM, RF and XGBoost), hyperparameter tuning, and model selection. Before models fitting, missing data were imputed using the K-Nearest Neighbours algorithm, and synthetic data generation was performed within each training fold using the R synthpop package.
Results
The literature search identified 35 studies meeting the inclusion criteria: 9 methodological papers and 26 empirical studies. The most used approaches for sample size determination for ML models are an a priori formula, an empirical learning curve method, and a parametric learning curve extrapolation. We applied them to the STRATEGY dataset, with the aim of predicting the modified Barthel Index (mBI) at discharge.
The three sample size determination approaches yielded concordant, but numerically different estimates. The a priori method produced the largest required sample size, consistent with its conservative analytical nature, followed by the empirical learning curve and the parametric extrapolation method. Notably, all three sample size estimates exceeded the sample size used by Campagnini et al. (2025, Scientific Reports) in a similar post-stroke rehabilitation population. Although not directly comparable, the lower predictive performance reported by Campagnini et al. is consistent with their smaller sample size, which was below estimates obtained in our post hoc assessment.
Conclusion
The sample size estimates obtained in this study should not be interpreted as fixed thresholds, but as complementary reference points that may support study planning. To the best of our knowledge, this is one of the first works to systematically map and compare the available methods for sample size determination in ML-based clinical prediction studies.
Downloads
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Sofia Partilora, Chiara Marzi, Valeriano Videtta , Silvia Pancani, Ilaria Pellegrini, Jorge Navarro, Lucia Falco, Caterina Tramonti, Ersilia Lucenteforte, Andrea Mannini, Francesca Cecchi

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
How to Cite
Published 2026-09-22


