TruaceTracing the truth around AIWednesday, August 26, 2026
The Index

What the evidence says.What the public feels.

Ranks distinct AI gain and problem claims from the published record. Scores reward impact, independent source strength, scale, confidence, and recency.

1,155 results
Show filters and sorting

AI gains · 641

84
GainHealth· Newly added· Evidence: Moderate (1 source)

An interpretable multi-instance learning method for accurate differentiation of malignant and benign laryngeal lesions in laryngoscopy: Results IMIL-Net achieved the highest diagnostic performance, with a mean area under the curve (AUC) of 0.975 (95% CI 0.959-0.991), accuracy of 0.915 (95% CI 0.883-0.947), sensitivity of 0.876 (95% CI 0.803-0.949), and specificity of 0.945 (95% CI 0.910-0.980).

Background Laryngeal cancer is a significant global health issue with high mortality, and early diagnosis is critical for survival. Developing accurate diagnostic models for laryngoscopy can reduce potential repeated biopsies and lessen the patient burden, representing an urgent clinical need.

Impact 30%69
Evidence 25%95
Scale 20%85
Confidence 15%87
Recency 10%99

Updated Aug 22, 2026 · TRV-2026-0850

84
GainHealth· Stable· Evidence: Moderate (1 source)

Systematic review of 35 studies found classical AI models for bipolar disorder achieved pooled AUC 0.80 for long-term maintenance response and 85%-97% accuracy for safety and dose optimization.

A PRISMA-guided systematic review of 35 studies examined classical AI for treatment optimization in adult bipolar disorder across five outcomes: acute response, long-term maintenance, relapse/readmission, safety/dose, and brain aging/phenotyping. By the July 2026 publication date, pooled performance ranged from modest for acute response (AUC 0.68) to moderate-to-high for maintenance (AUC 0.80) and high accuracy for safety/dose (85%-97%).

Impact 30%69
Evidence 25%95
Scale 20%85
Confidence 15%87
Recency 10%93

Updated Jul 22, 2026 · TRV-2026-0511

84
GainClimate· Stable· Evidence: Moderate (1 source)

Adopting best practices for AI servers in the USA could cut projected carbon emissions by up to 73% and water footprints by up to 86% between 2024 and 2030.

Published November 10 2025 in Nature Sustainability, the study models the sustainability implications of rapidly expanding generative AI server installations across the United States, projecting annual water and carbon footprints through 2030 and testing mitigation options.

Impact 30%69
Evidence 25%95
Scale 20%85
Confidence 15%87
Recency 10%93

Updated Jul 22, 2026 · TRV-2026-0474

84
GainHealth· Stable· Evidence: Moderate (1 source)

Multi-Scale Feature Fusion model combining U-Net segmentation, EfficientNet and attention autoencoder features fused via CCA and YOLO classification achieved 99.95% accuracy on apple leaf disease datasets, enabling early detection for sustainable agriculture.

Researchers described a Multi-Scale Feature Fusion system for apple leaf disease identification that segments diseased tissue with U-Net, cleans background with Rank Order Fuzzy filtering, extracts features with EfficientNet and an Attention-based Autoencoder, fuses them with Canonical Correlation Analysis, and classifies with YOLO. Tested on five apple leaf datasets, it reported 99.95% classification accuracy with cross-validation and statistical testing.

Impact 30%69
Evidence 25%95
Scale 20%85
Confidence 15%87
Recency 10%93

Updated Jul 19, 2026 · TRV-2026-0262

AI problems · 514

83
ProblemHealth· Newly added· Evidence: Moderate (1 source)

In the same tertiary GDD cohort, the Youden-optimal model missed 33.9% of children who progressed to ID and had NPV 31.7%, limiting safe rule-out and indicating poor transportability beyond high-prevalence referral settings.

Researchers retrospectively analyzed 2453 children diagnosed with GDD between January 2014 and December 2023 at a provincial tertiary children's rehabilitation centre, followed to at least 60 months. Using 28 predictors across perinatal, developmental, neuroimaging, electrophysiological, genetic and comorbidity domains, they trained L2- and L1-regularized logistic regression, random forest, XGBoost and LightGBM, with Platt scaling for the L2 model, and evaluated on a held-out 30% test set.

Impact 30%63
Evidence 25%95
Scale 20%85
Confidence 15%87
Recency 10%100

Updated Aug 25, 2026 · TRV-2026-0878

82
ProblemHealth· Stable· Evidence: Moderate (1 source)

Even the domain-specific M4CXR model was inconsistent with reference findings in 25.2% of chest radiograph cases and did not significantly change RADPEER discrepancy rates versus original interpretation.

In a retrospective study published July 11 2026, investigators tested 500 chest radiographs from one tertiary center with two AI systems, M4CXR and ChatGPT-4o, having four radiologists score AI-generated reports for finding detection and RADPEER discrepancies. M4CXR reached 55.8% complete concordance versus 19.8% for GPT-4o and reduced mean reporting time to 16.3 seconds from 179.2 seconds unaided.

Impact 30%63
Evidence 25%95
Scale 20%85
Confidence 15%87
Recency 10%93

Updated Jul 20, 2026 · TRV-2026-0312

82
ProblemHealth· Stable· Evidence: Moderate (1 source)

Fully automated cardiomegaly screening accuracy dropped to 76.07% on the external OpenI dataset, showing domain-shift vulnerability, while manual CTR measurement remains a clinical bottleneck and existing deep models suffer from algorithmic bloating.

A peer-reviewed study published July 16, 2026 describes UBNet-Seg, a lightweight 2.3-million-parameter U-Net variant that infers cardiomegaly from lung field geometry rather than explicit heart segmentation. Trained on 11,748 images, it was evaluated on external NIH and OpenI chest X-ray datasets, reporting 95.85% lung Dice and 0.05-second inference.

Impact 30%63
Evidence 25%95
Scale 20%85
Confidence 15%87
Recency 10%92

Updated Jul 18, 2026 · TRV-2026-0257

82
ProblemCrime· Stable· Evidence: Moderate (1 source)

7.2% of evaluated MCP servers contained general vulnerabilities and 5.5% exhibited MCP-specific tool poisoning, part of eight distinct vulnerabilities largely distinct from traditional software flaws.

In a first large-scale empirical study published May 2026, researchers examined 1,899 open-source Model Context Protocol servers, the standard introduced by Anthropic in late 2024 to unify tool calling for Foundation Models. Using health metrics and a combined general and MCP-specific scanner, they measured adoption signals and code quality across the ecosystem.

Impact 30%63
Evidence 25%95
Scale 20%85
Confidence 15%87
Recency 10%91

Updated Jul 13, 2026 · TRV-2026-0137

Recomputed live from the record · Aug 26, 2026, 10:54 PM