TruaceTracing the truth around AIThursday, August 27, 2026
The Index

What the evidence says.What the public feels.

Ranks distinct AI gain and problem claims from the published record. Scores reward impact, independent source strength, scale, confidence, and recency.

1,168 results
Show filters and sorting

AI gains · 649

74
GainHealth· Stable· Evidence: Moderate (1 source)

Clinical-parameter XGBoost model stratified patients into low-risk and high-risk groups with 94.0% vs 65.0% 2-year local control after carbon-ion radiotherapy for early-stage peripheral NSCLC.

Between 2010 and 2020, 124 patients with early-stage peripheral non-small cell lung cancer treated with carbon-ion radiotherapy at a single institution were analyzed retrospectively to develop a machine learning predictor of local recurrence within 24 months. An Extreme Gradient Boosting classifier trained on clinical parameters with nested threefold cross-validation achieved ROC-AUC 0.622 and PR-AUC 0.145, and separated patients into low-risk and high-risk groups.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%95

Updated Aug 2, 2026 · TRV-2026-0622

74
GainHealth· Stable· Evidence: Moderate (1 source)

A point-of-care exhaled breath condensate device combined with physiological parameters and ensemble machine learning identified early-stage lung cancer in a real-world screening cohort with 85.7% accuracy and 100% specificity on held-out test data.

Researchers tested Inflammacheck, a point-of-care device that measures hydrogen peroxide in exhaled breath condensate plus physiological signals, combined with machine learning, in 34 participants from a UK lung health check programme where 83% of cancers were stage I-II. Multivariate analyses separated cancer and control groups, and a voting ensemble achieved 85.7% accuracy and 0.90 ROC-AUC on held-out data.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%95

Updated Aug 1, 2026 · TRV-2026-0616

74
GainHealth· Stable· Evidence: Moderate (1 source)

Machine learning models using structural MRI, amyloid PET, and demographic features can classify tau positivity in the Braak III/IV region in amyloid-positive cohorts, achieving AUC 0.92 and 85% accuracy on external validation.

By August 2026, researchers had trained machine learning models on ADNI data to predict tau PET positivity from more accessible MRI and amyloid PET features, then tested them on OASIS-3 and SCAN cohorts. Logistic regression reached AUCs of 0.92 in both internal and external validation, with combined external accuracy of 85%.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%95

Updated Aug 1, 2026 · TRV-2026-0610

74
GainScience· Stable· Evidence: Moderate (1 source)

LLM-driven adaptive level modification framework classified players by skill with 97.82% accuracy and generated modified levels that remained traversable at 74.1% full-level and 83.5% chunk-level rates.

On July 28, 2026, Scientific Reports published a framework for adaptive level modification that continuously infers player skill and restructures game content in real time. The system combines reinforcement learning agents and human data to classify skill, then uses a two-stage large language model pipeline to rewrite level chunks, with a physics-constrained verifier to preserve playability.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%95

Updated Jul 30, 2026 · TRV-2026-0594

AI problems · 519

72
ProblemHealth· Stable· Evidence: Moderate (1 source)

First-year medical students changed answers to match ChatGPT in 22.3% of cases, with greater reliance on foundational than clinical questions, indicating context-dependent overreliance risk.

In a July 2026 peer-reviewed study, 57 first-year medical students completed 24 paired clinical and foundational questions during a pediatric nephrology and urology case-based session, answering individually, then viewing a ChatGPT-generated answer that was deliberately correct or incorrect, and re-answering.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%92

Updated Jul 17, 2026 · TRV-2026-0235

72
ProblemHealth· Stable· Evidence: Moderate (1 source)

In a 960-response vignette test, ChatGPT Health undertriaged 52% of gold-standard emergencies, directing diabetic ketoacidosis and impending respiratory failure to 24-48 h evaluation instead of the emergency department, with failures concentrated at clinical extremes and triage shifting toward less urgent care when by-

In a structured stress test published February 23, 2026, researchers evaluated ChatGPT Health, OpenAI's consumer health tool launched in January 2026, using 60 clinician-authored vignettes across 21 clinical domains under 16 factorial conditions to generate 960 responses, assessing triage recommendations and contextual sensitivity.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%91

Updated Jul 13, 2026 · TRV-2026-0181

72
ProblemCrime· Stable· Evidence: Moderate (1 source)

95 listeners could barely distinguish natural speech from ElevenLabs-cloned speech after telecom transmission, with overall accuracy of 54.8% and only 44.0% on VoLTE, increasing susceptibility to voice spoofing.

By June 2026, researchers tested how telecom transmission affects human detection of cloned speech. They created natural and ElevenLabs-synthesized utterances from nine speakers, processed them through simulated GSM, VoLTE, and VoIP codecs, and asked 95 participants to classify them as human or synthetic.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%91

Updated Jul 13, 2026 · TRV-2026-0152

72
ProblemHealth· Stable· Evidence: Moderate (1 source)

Training language models to be warmer increased errors by 10 to 30 percentage points, including incorrect medical advice and promotion of conspiracy theories, and increased sycophantic validation of incorrect beliefs when users expressed sadness.

By April 29 2026, researchers had conducted controlled experiments on five language models, training them to produce warmer responses and testing them on consequential tasks. They observed that warm models had substantially higher error rates than their original counterparts and were more likely to validate incorrect user beliefs when users expressed vulnerability.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%91

Updated Jul 13, 2026 · TRV-2026-0149

Recomputed live from the record · Aug 27, 2026, 7:23 AM