TruaceTracing the truth around AIWednesday, August 5, 2026
Health·The Trace·Automated dual reading·Published 2026-08-03

performance of federated learning models trained on patient data for health services research

Source article: Performance of federated learning models in health services research: A systematic review and meta-analysis

Purpose This study aimed to assess the performance of federated learning (FL) models and compare their performance with local and centralized models. Methods We conducted a systematic search of Ovid MEDLINE and PubMed from inception to June 10, 2025, to identify studies using patient data to train or validate FL algorithms and reporting at least one model performance outcome. Two reviewers independently screened articles and extracted data on study characteristics, FL frameworks and model training methodologies,…

TRV-2026-0627Peer-reviewedPermanent record — cite & verify
Trace impact reading

Contested: both sides are scored from claims and sources, not community votes.

P 74The P score combines the specificity and measured human impact of the grounded problem claim with the strength of this Trace’s cited sources.G 74The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.
Performance of federated learning models in health services research: A systematic review and meta-analysis

An assessment, survey, and systems engineering design of information sharing and discovery systems in a network-centric environment by De Soto, Kristine M.. Public domain

The quick read

A systematic review and meta-analysis of 13 studies covering 247 sites and 158,435 patient samples evaluated federated learning models, mostly using FedAvg, against local and centralized models on diagnostic performance metrics.

The findings matter because they quantify the privacy-performance trade-off for multi-site clinical AI, showing gains over isolated models but small losses versus pooled data, while highlighting that inconsistent reporting limits confidence in generalizing results to broader health research use.

Main points
  • Systematic review included 13 studies involving 247 sites and 158,435 samples, with 8 studies in meta-analysis.
  • Most FL models (85%) used the Federated Averaging (FedAvg) algorithm for parameter aggregation across sites.
  • Search covered Ovid MEDLINE and PubMed from inception to June 10, 2025 for studies using patient data to train or validate FL algorithms.
Gain

Federated learning models trained on distributed patient data improved predictive performance over single-site local models across AUC, F1, sensitivity, PPV and PRAUC.

Problem

Federated learning models showed modest performance losses compared with centralized models trained on pooled patient data across AUC, F1, sensitivity and PRAUC.

The rundown

Reviewers screened Ovid MEDLINE and PubMed to June 10, 2025, extracting study characteristics, FL frameworks, and performance metrics, following PRISMA extension for Diagnostic Test Accuracy Studies.

Within each machine learning task, FL metrics were summarized using medians and interquartile ranges and directly compared to local and centralized models trained on the same patient datasets.

What this doesn’t fix

Broader application is constrained by lack of standardized reporting and methodological transparency in current studies.

Sources

Reader signal

How should this claim be treated?

The debate