Sample Size Determination in Studies Involving Machine Learning Algorithms: A Case Study on Post-Stroke Rehabilitation

Authors

DOI:

https://doi.org/10.54103/2282-0930/32117

Abstract

Introduction

Sample size determination remains an important methodological challenge in the development of machine learning (ML) prediction models in clinical studies. Unlike classical statistical approaches such as regression-based methods, there is still no widely accepted procedure for estimating the amount of data required to develop reliable ML models, and current practice remains highly variable and largely based on informal or heterogeneous criteria.

 

Aims

This study had two main aims: 1) to identify and critically compare the methods proposed for sample size determination in ML-based clinical prediction studies through a systematic review; and 2) to apply and compare the methods identified in the systematic review within a real-world case study based on the STRATEGY study, a multicentre Italian cohort originally including 743 post-stroke patients admitted to intensive inpatient rehabilitation.

 

Methods

The systematic review was conducted using searches in three bibliographic databases: MEDLINE, the Cochrane Library, and EMBASE. In addition, a forward citation search was performed in Scopus to identify studies that applied the identified methodological approaches to justify sample size in their own ML-based analyses.

The identified methods were then applied to the STRATEGY dataset. Model development and evaluation were conducted using a nested validation framework. First, the dataset was split into a development set (80%) and an independent test set (20%). Within the development set, an inner 5-fold cross-validation procedure was used for models training (i.e. SVM, RF and XGBoost), hyperparameter tuning, and model selection. Before models fitting, missing data were imputed using the K-Nearest Neighbours algorithm, and synthetic data generation was performed within each training fold using the R synthpop package.

 

Results

The literature search identified 35 studies meeting the inclusion criteria: 9 methodological papers and 26 empirical studies. The most used approaches for sample size determination for ML models are an a priori formula, an empirical learning curve method, and a parametric learning curve extrapolation. We applied them to the STRATEGY dataset, with the aim of predicting the modified Barthel Index (mBI) at discharge.

The three sample size determination approaches yielded concordant, but numerically different estimates. The a priori method produced the largest required sample size, consistent with its conservative analytical nature, followed by the empirical learning curve and the parametric extrapolation method. Notably, all three sample size estimates exceeded the sample size used by Campagnini et al. (2025, Scientific Reports) in a similar post-stroke rehabilitation population. Although not directly comparable, the lower predictive performance reported by Campagnini et al. is consistent with their smaller sample size, which was below estimates obtained in our post hoc assessment.

 

Conclusion

The sample size estimates obtained in this study should not be interpreted as fixed thresholds, but as complementary reference points that may support study planning. To the best of our knowledge, this is one of the first works to systematically map and compare the available methods for sample size determination in ML-based clinical prediction studies.

Downloads

Download data is not yet available.

Downloads

Published

2026-09-22

How to Cite

1.
Sample Size Determination in Studies Involving Machine Learning Algorithms: A Case Study on Post-Stroke Rehabilitation. ebph [Internet]. 2026 Sep. 22 [cited 2026 Sep. 25]; Available from: https://riviste.unimi.it/index.php/ebph/article/view/32117
Received 2026-06-29
Published 2026-09-22