← PublicationsJournal article
2026
Can one model fit all? Evaluating foundation models for time series forecasting across clinical medicine
Gernot Pucher, Amin Dada, Aurel Agbodoyetin, Felix Nensa, Martin Schuler, Hans Christian Reinhardt, Jens Kleesiek, Christopher Sauer
Artificial Intelligence in Medicine
Abstract
Artificial intelligence (AI) is increasingly integrated into clinical medicine, with foundation models emerging as an alternative to task-specific models for forecasting longitudinal healthcare data. These models, pre-trained on large datasets, promise broad applicability across clinical domains, yet their real-world performance and generalizability remain underexplored. To address this gap, we evaluated foundation and task-specific models across diverse clinical use cases, focusing on zero-shot performance, cross-hospital transportability, the impact of fine-tuning, and potential clinical implications. We used data from University Hospital Essen, Germany, two nearby regional hospitals, and the MIMIC-IV database to define six clinical time series use cases, including forecasting of vital signs, laboratory values, and hospital capacity. Transformer-based foundation models were compared in zero-shot and fine-tuned settings to task-specific approaches, including neural networks, gradient boosting, AutoML ensembles, and statistical models. We also assessed predictive value for guiding treatment decisions by dichotomizing forecasts. Zero-shot foundation models frequently approached the performance of optimized task-specific models. Fine-tuning further improved performance, with Chronos and TimesFM ranking among the best-performing models 19 and 18 times, respectively, compared to 21 times for AutoML ensembles. Foundation models showed superior transportability across hospital settings and patient populations. However, variations in forecasting strategies influenced their positive and negative predictive values in clinical decision-making contexts. These results suggest that foundation models are viable for clinical time series forecasting, particularly where generalizability is crucial. Their flexibility and zero-shot capabilities reduce the need for retraining, potentially lowering barriers to adoption and challenging the role of domain-specific models in clinical practice.
publications