performance of federated learning models trained on patient data for health services research
Source article: Performance of federated learning models in health services research: A systematic review and meta-analysis
Purpose This study aimed to assess the performance of federated learning (FL) models and compare their performance with local and centralized models. Methods We conducted a systematic search of Ovid MEDLINE and PubMed from inception to June 10, 2025, to identify studies using patient data to train or validate FL algorithms and reporting at least one model performance outcome. Two reviewers independently screened articles and extracted data on study characteristics, FL frameworks and model training methodologies,…
Contested: both sides are scored from claims and sources, not community votes.
An assessment, survey, and systems engineering design of information sharing and discovery systems in a network-centric environment by De Soto, Kristine M.. Public domain
A systematic review and meta-analysis of 13 studies covering 247 sites and 158,435 patient samples evaluated federated learning models, mostly using FedAvg, against local and centralized models on diagnostic performance metrics.
The findings matter because they quantify the privacy-performance trade-off for multi-site clinical AI, showing gains over isolated models but small losses versus pooled data, while highlighting that inconsistent reporting limits confidence in generalizing results to broader health research use.
- Systematic review included 13 studies involving 247 sites and 158,435 samples, with 8 studies in meta-analysis.
- Most FL models (85%) used the Federated Averaging (FedAvg) algorithm for parameter aggregation across sites.
- Search covered Ovid MEDLINE and PubMed from inception to June 10, 2025 for studies using patient data to train or validate FL algorithms.
Federated learning models trained on distributed patient data improved predictive performance over single-site local models across AUC, F1, sensitivity, PPV and PRAUC.
Federated learning models showed modest performance losses compared with centralized models trained on pooled patient data across AUC, F1, sensitivity and PRAUC.
The rundown
Reviewers screened Ovid MEDLINE and PubMed to June 10, 2025, extracting study characteristics, FL frameworks, and performance metrics, following PRISMA extension for Diagnostic Test Accuracy Studies.
Within each machine learning task, FL metrics were summarized using medians and interquartile ranges and directly compared to local and centralized models trained on the same patient datasets.
Broader application is constrained by lack of standardized reporting and methodological transparency in current studies.
Sources
- Peer-reviewedAdvances in Medical Sciences2026-08-01
How should this claim be treated?
ace
The debate