TruaceTracing the truth around AIWednesday, August 26, 2026
The Index

What the evidence says.What the public feels.

Ranks distinct AI gain and problem claims from the published record. Scores reward impact, independent source strength, scale, confidence, and recency.

1,155 results
Show filters and sorting

AI gains · 641

78
GainPolicy· Stable· Evidence: Moderate (1 source)

Chinese and South Korean regulatory toolkits for AI journalism on platforms like Toutiao and Naver were found to mitigate risks of digital infodemics and algorithmic bias.

This comparative study analyzed China and South Korea's distinct approaches to governing AI journalism and algorithmic news curation, examining policy documents and evidence from Toutiao and Naver to assess how each balances fairness and accountability.

Impact 30%49
Evidence 25%95
Scale 20%85
Confidence 15%87
Recency 10%92

Updated Jul 17, 2026 · TRV-2026-0247

78
GainHealth· Stable· Evidence: Moderate (1 source)

AI combined with multimodal retinal imaging provides a non-invasive method for early risk stratification and screening for cardiovascular, metabolic and neurodegenerative disorders using retinal vasculature and nerve layer changes.

A May 2026 review in Graefe's Archive describes AI combined with multimodal retinal imaging as a non-invasive approach to detect and monitor systemic vascular and neurodegenerative conditions. It outlines how fundus photography, OCT, OCTA and metabolic-sensitive imaging capture retinal vascular and nerve changes that reflect cardiovascular, metabolic and neurological disease, analyzed with deep learning and multimodal fusion.

Impact 30%49
Evidence 25%95
Scale 20%85
Confidence 15%87
Recency 10%92

Updated Jul 13, 2026 · TRV-2026-0192

78
GainHealth· Stable· Evidence: Moderate (1 source)

AI systems in the operating room using multimodal data from patients, teams, robots and environment to provide situational awareness and intraoperative decision-making that optimizes surgical actions.

This peer-reviewed analysis from May 2026 examines how AI and robotics ecosystems are entering the operating room, using multimodal data from patients, staff, robots and the environment for workflow recognition, performance benchmarking and decision support, while robots evolve toward autonomous systems with human-in-the-loop control.

Impact 30%49
Evidence 25%95
Scale 20%85
Confidence 15%87
Recency 10%92

Updated Jul 13, 2026 · TRV-2026-0191

78
GainHealth· Stable· Evidence: Moderate (1 source)

National EHR networks covering up to more than 200 million patients can support learning health systems by enabling large-scale aggregation and benchmarking for ML/AI development.

Researchers conducted an environmental scan through September 2025 of 23 US national EHR networks that aggregate patient-level data, ranging from under 1 million to over 200 million patients, and reviewed 34 ML/AI studies built on them. Most networks used common data models, yet few models were prospectively evaluated or integrated into clinical workflows.

Impact 30%49
Evidence 25%95
Scale 20%85
Confidence 15%87
Recency 10%91

Updated Jul 13, 2026 · TRV-2026-0187

AI problems · 514

74
ProblemClimate· Newly added· Evidence: Moderate (1 source)

Fixed 20-epoch training without checkpoint selection wasted most compute, with 78% to 84% of total emissions occurring after the optimal checkpoint had already been reached.

In a study published August 12, 2026, researchers quantified CO2eq emissions for training ResNet-50, DenseNet-121 and EfficientNet-B0 on 128,907 chest radiographs for 20 epochs. They found validation loss minima at median epochs 2 to 4, meaning most emissions occurred after the best checkpoint, and compared retrospective selection, prospective early stopping, and fixed-epoch training on AUC and energy use.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%98

Updated Aug 14, 2026 · TRV-2026-0753

74
ProblemHealth· Stable· Evidence: Moderate (1 source)

Item Response Theory produced an unacceptably low 68% sensitivity for the total score at threshold 24, and logistic regression yielded only 16%-60% sensitivity for ADHD status in the same driver sample.

A 2026 peer-reviewed study validated the Persian Conners' Adult ADHD Rating Scale Short Version in 298 male taxi drivers in Iran, mean age 36.8, to establish occupational screening thresholds. Using a 198/100 train-test split, the authors compared ROC, item response theory, logistic regression and Random Forest approaches for cutoff selection.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%95

Updated Aug 3, 2026 · TRV-2026-0633

74
ProblemHealth· Stable· Evidence: Moderate (1 source)

Across 29 standardized clinical vignettes, all 21 tested LLMs failed differential diagnosis in over 80% of cases, indicating they have not achieved the reasoning needed for safe clinical deployment.

Researchers evaluated 21 off-the-shelf large language models, including GPT-5, Claude 4.5 Opus, Gemini 3.0 and Grok 4, on 29 standardized MSD Manual clinical vignettes representing 16,254 responses scored by medical students. Using the PrIME-LLM composite across differential diagnosis, diagnostic testing, final diagnosis, management, and miscellaneous reasoning, scores ranged from 0.64 to 0.78.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%91

Updated Jul 13, 2026 · TRV-2026-0145

73
ProblemHealth· Newly added· Evidence: Moderate (1 source)

Chi-squared selection with a neural network achieved the highest mean accuracy (74.78%; AUC 0.77).

Nottingham histological grading is central to breast cancer prognosis and treatment planning, but conventional pathological assessment is labor-intensive and subject to inter-observer variability. Radiomics and machine learning may support noninvasive preoperative grade prediction.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%99

Updated Aug 23, 2026 · TRV-2026-0858

Recomputed live from the record · Aug 27, 2026, 3:52 AM