TruaceTracing the truth around AIFriday, August 28, 2026
The Index

What the evidence says.What the public feels.

Ranks distinct AI gain and problem claims from the published record. Scores reward impact, independent source strength, scale, confidence, and recency.

1,182 results
Show filters and sorting

AI gains · 658

68
GainHealth· Stable· Evidence: Moderate (1 source)

An exploratory Light Gradient Boosting Machine model predicted 1-week willingness to reuse rubber dam isolation after microscopic root canal treatment with AUC 0.939 in held-out test and 0.983 in temporal validation.

Researchers retrospectively analyzed 306 patients who had microscopic root canal treatment with rubber dam isolation between May and November 2025, defining willingness to reuse at 1-week follow-up as the outcome, with 246 willing and 60 unwilling. They trained six models on 26 variables and found the LightGBM model retained 12 predictors and achieved the highest exploratory AUCs of 0.939 and 0.983.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%95

Updated Aug 3, 2026 · TRV-2026-0631

68
GainHealth· Stable· Evidence: Moderate (1 source)

LLM-based staged extraction framework localized PVC origin as left versus right from 12-lead ECG images with discrimination comparable to a CNN baseline while providing a traceable stepwise diagnostic process.

Researchers tested whether large language models could interpret 12-lead ECG images to distinguish left- versus right-sided origins of premature ventricular contractions in 157 patients who had undergone successful catheter ablation. By August 2026 they reported a staged extraction framework that produced a traceable stepwise process and a continuous score whose discrimination was numerically similar to a CNN baseline.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%95

Updated Aug 3, 2026 · TRV-2026-0629

68
GainHealth· Stable· Evidence: Moderate (1 source)

Researchers developed a nomogram using AI-derived CCTA measures of pericoronary adipose tissue and plaque plus HbA1c to predict progression of non-obstructive coronary lesions in T2DM patients.

Between 2019 and 2024, researchers retrospectively followed 114 patients with type 2 diabetes and non-obstructive coronary artery disease who had baseline CCTA. Using AI-derived measurements of pericoronary adipose tissue and coronary plaques combined with clinical labs, they built a logistic regression nomogram to distinguish 48 patients who later had infarction, revascularization, or stenosis 265 50% from 66 who did not.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%95

Updated Aug 3, 2026 · TRV-2026-0628

68
GainHealth· Stable· Evidence: Moderate (1 source)

Federated learning models trained on distributed patient data improved predictive performance over single-site local models across AUC, F1, sensitivity, PPV and PRAUC.

A systematic review and meta-analysis of 13 studies covering 247 sites and 158,435 patient samples evaluated federated learning models, mostly using FedAvg, against local and centralized models on diagnostic performance metrics.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%95

Updated Aug 3, 2026 · TRV-2026-0627

AI problems · 524

68
ProblemHealth· Stable· Evidence: Moderate (1 source)

As of the June 2025 search cutoff, evidence did not support claims that AI delivers precision medical education through OSCEs, with weaker performance in relational, situated, and culturally mediated competencies.

By July 2026, this scoping review had mapped the literature on AI in OSCEs, screening 421 records and including 22 studies across health professions education. It categorized uses in preparation, station construction, scoring, and delivery, and stratified findings by evidence maturity using FACETS, SAMR, and P4 frameworks.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Jul 20, 2026 · TRV-2026-0325

68
ProblemHealth· Stable· Evidence: Moderate (1 source)

MMAI lost independent prognostic value after adjustment for CAPRA clinical risk score and did not outperform established clinical tools for postprostatectomy outcome prediction.

By July 2026, investigators had tested a multimodal AI model originally validated in prostate biopsy specimens on a tissue microarray of 424 prostatectomy cases, generating scores for 414 patients. At 10 years, recurrence-free survival was 74% and metastasis-free survival 96%, and MMAI scores were associated with both endpoints in univariable models.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Jul 20, 2026 · TRV-2026-0324

68
ProblemHealth· Stable· Evidence: Moderate (1 source)

At the low-volume tertiary care hospital, implementing PhenoMATRIX alone without a timely release workflow was associated with increased time to result reporting due to delays between result availability and reporting.

By July 10 2026, a dual-center Canadian study reported before-and-after results for PhenoMATRIX, an AI-based software that provides continuous culture sorting and interpretation support for urine cultures on laboratory automation. Both a low-volume tertiary hospital and a high-volume community lab saw earlier availability of interpretable results, with measured TTRR changes of approximately 1.3 hours with PM+ automated release and approximately 5.3 hours with earlier manual review.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Jul 20, 2026 · TRV-2026-0322

68
ProblemHealth· Stable· Evidence: Moderate (1 source)

GPT-4o showed optimism and positional biases in risk-of-bias assessment because it relied on probabilistic language patterns rather than structured clinical reasoning.

By December 2025, researchers conducted a systematic review and meta-analysis of 13 trials (n=684) of five emerging biologics for osteogenesis imperfecta, using GPT-4o to perform parallel title/abstract and full-text screening and to assist risk-of-bias assessment. The AI workflow achieved 97.4% sensitivity at abstract level and 88.9% at full-text, reducing total screening time by over 95% with substantial agreement to humans (kappa 0.778).

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Jul 20, 2026 · TRV-2026-0321

Recomputed live from the record · Aug 28, 2026, 10:47 AM