TruaceTracing the truth around AIThursday, August 27, 2026
The Index

What the evidence says.What the public feels.

Ranks distinct AI gain and problem claims from the published record. Scores reward impact, independent source strength, scale, confidence, and recency.

1,169 results
Show filters and sorting

AI gains · 649

68
GainHealth· Newly added· Evidence: Moderate (1 source)

AI models predicted operative difficulty in laparoscopic cholecystectomy with pooled discrimination of 0.848 in training and 0.818 in validation, with ensemble and multimodal models performing best.

By August 13, 2026, a systematic review and meta-analysis of 18 studies found AI models predicted laparoscopic cholecystectomy difficulty with pooled AUCs of 0.848 in training and 0.818 in validation, with ensemble models reaching 0.889 and 0.861. The review searched four databases to March 2, 2026 and used PROBAST and GRADE to assess bias and certainty.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%98

Updated Aug 16, 2026 · TRV-2026-0787

68
GainHealth· Newly added· Evidence: Moderate (1 source)

Causal AI integrated into healthcare informatics systems improves personalized clinical decision-making by estimating individual treatment effects and simulating intervention outcomes.

On 2026-08-13, a review in Personalized Medicine examined integration of causal artificial intelligence and data-driven decision intelligence within healthcare informatics to advance personalized medicine. Using a narrative review of literature from PubMed, Scopus, Web of Science, IEEE Xplore and ScienceDirect, the authors synthesized evidence on causal inference methods and clinical applications.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%98

Updated Aug 16, 2026 · TRV-2026-0786

68
GainLabor· Newly added· Evidence: Moderate (1 source)

A gradient boosting machine trained on survey data from 318 ICU nurses predicted high moral distress with high discrimination using six predictors, preserving accuracy after feature reduction.

In a multicentre cross-sectional study of 318 ICU nurses in China, researchers developed a machine learning risk-profiling model for moral distress, which was present in 28.6% of participants. A gradient boosting machine achieved the best balanced performance and, after SHAP-guided reduction, retained full accuracy with six predictors: monthly night shifts, financial responsibility role, psychological resilience, sleep quality, nurse-to-patient ratio, and weekly working hours.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%98

Updated Aug 16, 2026 · TRV-2026-0785

68
GainHealth· Newly added· Evidence: Moderate (1 source)

L2-regularized logistic regression trained on routine non-invasive paraclinical markers achieved stable discrimination with minimal generalization gap and robust calibration for early prediction of bacterial infections in infants aged 1 to 90 days.

Researchers retrospectively analyzed 306 infants aged 1 to 90 days hospitalized between 2014 and 2022 in Khorasan Razavi, Iran, using CSF culture via lumbar puncture as the gold standard, to train nine machine learning classifiers on routine non-invasive paraclinical markers with nested cross-validation and SHAP interpretation.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%98

Updated Aug 16, 2026 · TRV-2026-0783

AI problems · 520

68
ProblemScience· Stable· Evidence: Moderate (1 source)

Frontier Large Reasoning Models face a complete accuracy collapse beyond certain puzzle complexities and exhibit a counterintuitive scaling limit where reasoning effort declines despite adequate token budget.

Published September 23, 2025, this peer-reviewed study systematically tested frontier Large Reasoning Models that generate detailed thinking processes before answering. Using controllable puzzle environments to vary compositional complexity, the authors analyzed final accuracy and internal reasoning traces and compared LRMs to standard LLMs under equivalent inference compute.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Jul 22, 2026 · TRV-2026-0487

68
ProblemHealth· Stable· Evidence: Moderate (1 source)

Integration into routine care is constrained by limited explainability, data bias, lack of prospective trials, regulatory hurdles, and mixed real-world outcome evidence for decision support tools.

A September 2025 peer-reviewed review in Clinics and Practice synthesized 150 studies of AI in clinical medicine after screening 2047 PubMed records. It found strong diagnostic imaging performance with expert-level cancer detection, promise for CDSS in predicting sepsis and atrial fibrillation, and advances in surgical guidance, pathology diagnosis, and drug discovery via protein structure prediction.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Jul 22, 2026 · TRV-2026-0486

68
ProblemCrime· Stable· Evidence: Moderate (1 source)

Generative AI enables fraud attacks that erode Zero-Trust Architecture by using synthetic identities and context manipulation to increase false-negative rates, extend dwell time, bypass policies, and evade audit trails.

In a peer-reviewed survey published October 15, 2025, researchers analyzed 10 recent Zero-Trust Architecture surveys and 136 primary studies from 2022-2024 and found most controls lacked real-world validation. They argue generative AI attacks exploit those gaps and propose a seven-stage Cyber Fraud Kill Chain that maps synthetic identities, context manipulation, and adversarial telemetry to NIST SP 800-207 components.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Jul 22, 2026 · TRV-2026-0483

68
ProblemLifestyle· Stable· Evidence: Moderate (1 source)

Popular AI chatbots exhibit class-based regularities in how they portray the lifestyle and tastes of fictional personas across different occupations.

In a peer-reviewed study published October 2025, researchers conducted 39 interviews with ChatGPT, Gemini and Replika, prompting each to impersonate people in six occupational groups ranging from highly skilled professionals and humanities professors to blue-collar workers, construction workers, computer scientists and hairdressers. The qualitative analysis identified regularities in how the chatbots described everyday tastes and lifestyles that aligned with class distinctions.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%93

Updated Jul 22, 2026 · TRV-2026-0482

Recomputed live from the record · Aug 28, 2026, 3:22 AM