TruaceTracing the truth around AIThursday, August 27, 2026
The Index

What the evidence says.What the public feels.

Ranks distinct AI gain and problem claims from the published record. Scores reward impact, independent source strength, scale, confidence, and recency.

1,168 results
Show filters and sorting

AI gains · 649

74
GainHealth· Stable· Evidence: Moderate (1 source)

A microorganism-based random forest model built from plasma metagenomic profiles predicted subsequent infection in newly diagnosed hematological patients with AUC 0.942, identifying 99.1% of those who later developed infections, and improved to AUC 0.953 when combined with clinical metrics, supporting targeted prophyl-

In a prospective study registered as ChiCTR2100042992, investigators collected plasma for metagenomic next-generation sequencing from 230 newly diagnosed hematological patients before and after chemotherapy and used machine learning to map a complex microecological landscape linked to neutropenia and subsequent infection.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%96

Updated Aug 7, 2026 · TRV-2026-0672

74
GainHealth· Stable· Evidence: Moderate (1 source)

Manus architecture using raw 3D CBCT data detected and correctly diagnosed 95% of jaw lesions in 97 patients, outperforming 2D panoramic inputs.

A cross-sectional study tested four AI chatbots on 97 anonymized CBCT cases of jaw lesions, comparing performance on reconstructed 2D panoramic views and, for Manus, raw 3D DICOM data. Reports were scored for accuracy, relevance and feasibility, revealing statistically significant differences between systems.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%96

Updated Aug 5, 2026 · TRV-2026-0653

74
GainHealth· Stable· Evidence: Moderate (1 source)

In 298 Iranian male taxi drivers, ROC and Random Forest analysis of the Persian CAARS-S:SV identified total-score cutoffs with 88% sensitivity and 86.7% specificity for adult ADHD screening.

A 2026 peer-reviewed study validated the Persian Conners' Adult ADHD Rating Scale Short Version in 298 male taxi drivers in Iran, mean age 36.8, to establish occupational screening thresholds. Using a 198/100 train-test split, the authors compared ROC, item response theory, logistic regression and Random Forest approaches for cutoff selection.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%95

Updated Aug 3, 2026 · TRV-2026-0633

74
GainHealth· Stable· Evidence: Moderate (1 source)

A hybrid Mask R-CNN and YOLOv11 segmentation pipeline automated postoperative Pink Esthetic Score attribute assessment from intraoral photographs with over 82% accuracy per attribute and 79.3% total-score agreement within one point of experts.

Researchers developed and internally validated an anatomy-driven AI system to automate postoperative Pink Esthetic Score evaluation from intraoral photographs, using Mask R-CNN for tooth crowns and YOLOv11 for gingiva to derive measurements rather than end-to-end prediction. Tested against independent expert scoring of 82 photographs, the system reached 91.5% accuracy for mesial papilla and 82.9% to 86.6% for other attributes, with 57.3% exact total-score agreement.

Impact 30%69
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%95

Updated Aug 3, 2026 · TRV-2026-0626

AI problems · 519

72
ProblemHealth· Stable· Evidence: Moderate (1 source)

LLM summarisation of consultations showed a 1.47% hallucination rate and 3.45% omission rate, creating fidelity gaps that could compromise patient safety.

On 2025-05-13, a peer-reviewed framework was described for evaluating LLMs that automate summarising consultations into clinical notes. It combines an error taxonomy, iterative experimental comparisons, a clinical safety harm assessment, and the CREOLA interface, tested across 18 configurations with 12,999 clinician-annotated sentences.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Jul 24, 2026 · TRV-2026-0522

72
ProblemBusiness· Stable· Evidence: Moderate (1 source)

Those TFP gains are likely exaggerated and even more modest, predicted to be less than 0.53% over 10 years, because future AI effects will involve hard-to-learn tasks with many context-dependent factors and no objective outcome measures.

This peer-reviewed paper models AI's macroeconomic impact as task-level automation and complementarity, using Hulten's theorem to translate the fraction of tasks impacted and average cost savings into GDP and TFP effects. Using existing exposure estimates, it calculates no more than a 0.66% TFP increase over 10 years, then revises down to less than 0.53% after accounting for the shift from easy-to-learn to hard-to-learn tasks.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Jul 20, 2026 · TRV-2026-0378

72
ProblemHealth· Stable· Evidence: Moderate (1 source)

Among the same respondents, 70.8% cited technical reliability and 68.2% cited data privacy as top concerns about AI application for cervical screening, while 41.2% reported anxiety during result waiting and 58.1% struggled with medical terminology.

On July 10 2026, a peer-reviewed cross-sectional study reported results from 308 online questionnaire responses about cervical HPV screening experiences. Most respondents were urban women aged 25-35, 76.30% reported a history of HPV infection, and 91.56% had undergone TCT. The study measured current distress points and attitudes toward AI-assisted diagnosis.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Jul 20, 2026 · TRV-2026-0329

72
ProblemHealth· Stable· Evidence: Moderate (1 source)

The same e-CTA tool showed progressively lower sensitivity for more distal occlusions, dropping to 73% for distal M1 with only moderate agreement with experts, and excluded non-target occlusions from primary analysis, requiring adjunctive rather than standalone use.

Between May 2023 and May 2025, researchers retrospectively evaluated 531 multiphase CTA examinations from consecutive patients with suspected acute ischemic stroke at a single center, comparing Brainomix e-CTA automated LVO detection to expert neuroradiologist interpretation.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Jul 20, 2026 · TRV-2026-0307

Recomputed live from the record · Aug 27, 2026, 6:22 AM