TruaceTracing the truth around AIThursday, August 27, 2026
The Index

What the evidence says.What the public feels.

Ranks distinct AI gain and problem claims from the published record. Scores reward impact, independent source strength, scale, confidence, and recency.

1,168 results
Show filters and sorting

AI gains · 649

72
GainClimate· Newly added· Evidence: Moderate (1 source)

Prospective early stopping with patience 10 preserved chest radiograph classification performance while lowering total training emissions by up to 38% and raising carbon efficiency by up to 76% compared to fixed 20-epoch training.

In a study published August 12, 2026, researchers quantified CO2eq emissions for training ResNet-50, DenseNet-121 and EfficientNet-B0 on 128,907 chest radiographs for 20 epochs. They found validation loss minima at median epochs 2 to 4, meaning most emissions occurred after the best checkpoint, and compared retrospective selection, prospective early stopping, and fixed-epoch training on AUC and energy use.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%97

Updated Aug 14, 2026 · TRV-2026-0753

72
GainHealth· Newly added· Evidence: Moderate (1 source)

Large language models directed patients with musculoskeletal complaints to currently practicing, specialty-appropriate providers in the requested city, with ChatGPT achieving 100% appropriateness in the tested queries.

Researchers prompted ChatGPT, DeepSeek, and Gemini with standardized musculoskeletal complaints for Lynchburg, VA and Trumbull, CT, and judged whether recommended physicians were currently practicing locally in the relevant specialty and whether phone numbers were correct. By the August 13, 2026 publication date, ChatGPT was appropriate in all 17 recommendations, while Gemini and DeepSeek were appropriate in 43% and 40% of recommendations respectively.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%97

Updated Aug 14, 2026 · TRV-2026-0752

72
GainHealth· Stable· Evidence: Moderate (1 source)

YOLOv8x-pose and YOLOv8x-seg models automated landmark detection and apical segmentation for the Cameriere European method in children aged 5-13, achieving high detection accuracy and low measurement error in a first-stage validation.

This first-stage retrospective validation evaluated YOLOv8-based models to automate the anatomical inputs for the Cameriere European dental age estimation method using 4,050 panoramic radiographs of children aged 5-13. A YOLOv8x-pose model detected open-apex landmarks and a YOLOv8x-seg model segmented closed apices, with performance compared to manual reference annotations from CranioCatch software.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%96

Updated Aug 9, 2026 · TRV-2026-0715

72
GainHealth· Stable· Evidence: Moderate (1 source)

ChatGPT-5.0 achieved 75.77% accuracy and the highest sensitivity at 76.68% for orthodontic extraction decisions, performing comparably to XGBoost and significantly better than random forest, SVM, logistic regression and MLP.

A comparative study published August 7, 2026 evaluated ChatGPT-5.0 against five supervised machine learning algorithms for orthodontic extraction decisions. Using 520 cases (42.88% extraction, 57.12% non-extraction) and 23 clinical, cephalometric and photographic variables, with expert consensus as reference, ChatGPT-5.0 achieved 75.77% accuracy and 76.68% sensitivity under 5-fold cross-validation, compared to 78.08% accuracy for XGBoost.

Impact 30%63
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%96

Updated Aug 8, 2026 · TRV-2026-0689

AI problems · 519

68
ProblemHealth· Newly added· Evidence: Moderate (1 source)

Oncology AI systems can reproduce or amplify existing disparities across patient populations, and efforts to enforce fairness definitions often conflict with overall predictive performance.

As of the August 2026 commentary, AI was increasingly integrated into oncology for detection, risk stratification, treatment planning, and documentation. The authors reviewed evidence that these systems can reproduce or amplify disparities and examined technical sources of bias and competing statistical definitions of fairness.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%98

Updated Aug 18, 2026 · TRV-2026-0829

68
ProblemHealth· Newly added· Evidence: Moderate (1 source)

When used for prostate MRI reporting, LLMs can hallucinate measurements, flip negations, misstate laterality, and overstate cancer likelihood, creating patient-safety and accountability risks especially if reports are copied outside clinical governance.

Published August 17 2026 in Abdominal Radiology, this Perspective examines large language models applied to prostate MRI reporting, a task where laterality, sector, size, PI-RADS, and staging language directly affect biopsy and treatment decisions and where patients often see reports via portals before clinician discussion.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%98

Updated Aug 18, 2026 · TRV-2026-0828

68
ProblemHealth· Newly added· Evidence: Moderate (1 source)

Current AI systems for psychological assessment and treatment have not shown ability to produce meaningful and sustained clinical change, falling short due to memory limits, sycophancy, and focus on short-term helpfulness.

On 2026-08-17, a peer-reviewed framework paper argued that while large language models could augment psychological assessment and treatment, current technologies have not demonstrated sustained clinical benefit. The authors attribute this to poor integration of clinical science and to a duration mismatch between brief AI chats and months-long evidence-based treatments.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%98

Updated Aug 18, 2026 · TRV-2026-0821

68
ProblemLifestyle· Newly added· Evidence: Moderate (1 source)

Perceived risk around AIGC reduced fashion designers' feelings of autonomy, competence and relatedness, undermining psychological conditions for adoption.

Researchers examined why fashion designers adopt Artificial Intelligence Generated Content, which is described as increasingly used in creative design. Using the Stimulus-Organism-Response framework combined with Self-Determination Theory, they surveyed 318 Chinese fashion-design practitioners and analyzed 21 items with PLS-SEM to link perceived risk, social influence and facilitating conditions to autonomy, competence, relatedness and behavioral intention.

Impact 30%49
Evidence 25%95
Scale 20%35
Confidence 15%87
Recency 10%98

Updated Aug 18, 2026 · TRV-2026-0818

Recomputed live from the record · Aug 27, 2026, 1:36 PM