TruaceTracing the truth around AIWednesday, August 26, 2026

All traces

Skyer: a novel benchmark for evaluating the effectiveness of large language models in emergency department triage
HealthContested · G 67 / P 64

large language model triage performance in 55 pediatric ED scenarios measured by weighted accuracy accounting for over-triage and under-triage

Source article: Skyer: a novel benchmark for evaluating the effectiveness of large language models in emergency department triage

Problem

Despite higher scores, the evaluated large language models have identified limitations that prohibit them from replacing human experts for triage in overcrowded emergency departments.

Canadian Journal of Emergency Medicine
Gain

In 55 realistic pediatric scenarios evaluated by the Skyer benchmark, ChatGPT-4.5-preview and Gemini-2.5_05-06 achieved higher triage accuracy and weighted scores than human triage experts, with acceptable consistency across repeats.

Canadian Journal of Emergency Medicine
Performance evaluation of five major large language models in tuberculosis Q&A systems: A multidimensional assessment of readability, quality, and reliability
HealthContested · G 70 / P 74

LLM-generated pulmonary tuberculosis Q&A responses evaluated for patient-education quality and readability

Source article: Performance evaluation of five major large language models in tuberculosis Q&A systems: A multidimensional assessment of readability, quality, and reliability

Problem

The same five LLMs showed significant differences on several readability indices for TB education texts, creating uneven reading difficulty that may undermine patient understanding and adherence.

Frontiers in Public Health
Gain

In a 20-question TB Q&A test generating 100 responses, GPT-5 produced the most suitable patient-education texts as measured by C-PEMAT-P among five LLMs.

DIGITAL HEALTH
Mapping artificial intelligence integration in objective structured clinical examinations: A scoping review
HealthContested · G 75 / P 73

AI integration in objective structured clinical examinations for precision medical education

Source article: Mapping artificial intelligence integration in objective structured clinical examinations: A scoping review

Problem

As of the June 2025 search cutoff, evidence did not support claims that AI delivers precision medical education through OSCEs, with weaker performance in relational, situated, and culturally mediated competencies.

Medical Teacher
Gain

In health professions OSCEs, AI applications improved grading efficiency, feedback speed, and consistency for structured observable tasks, with personalization as the dominant P4 precision education alignment observed by June 2025.

Medical Teacher
Combining pathology artificial intelligence and genomic biomarkers to refine long-term postprostatectomy outcome prediction
HealthNegative state · G 68 / P 73

prediction of metastasis after prostatectomy using MMAI

Source article: Combining pathology artificial intelligence and genomic biomarkers to refine long-term postprostatectomy outcome prediction

Problem

MMAI lost independent prognostic value after adjustment for CAPRA clinical risk score and did not outperform established clinical tools for postprostatectomy outcome prediction.

JNCI: Journal of the National Cancer Institute
Gain

In 414 prostatectomy cases, the MMAI score derived from digitized pathology images was associated with long-term biochemical recurrence and metastasis, and when combined with genomic CCP score reached the highest discrimination for metastasis.

JNCI: Journal of the National Cancer Institute
Revealed: landmark Scottish AI project has no prospect of meeting renewables promise
OtherPositive state · G 62 / P 55

whether the Lanarkshire AI datacentre complex can be powered entirely by new on-site renewable energy as promised

Source article: Revealed: landmark Scottish AI project has no prospect of meeting renewables promise

Problem

The Lanarkshire AI datacentre complex has no prospect of meeting its promise of 1GW on-site renewable power, faces an acknowledged power provision issue, and will need to connect to the grid instead of operating behind-the-meter.

The Guardian
Gain

The Lanarkshire AI datacentre complex was presented as delivering jobs and prosperity while being powered entirely from on-site renewables with up to 1GW of new energy infrastructure by 2030.

The Guardian
Improving turnaround times with artificial intelligence in microbiology
HealthContested · G 73 / P 69

time to result reporting for urine cultures after AI-based PhenoMATRIX implementation in Canadian diagnostic laboratories

Source article: Improving turnaround times with artificial intelligence in microbiology

Problem

At the low-volume tertiary care hospital, implementing PhenoMATRIX alone without a timely release workflow was associated with increased time to result reporting due to delays between result availability and reporting.

Journal of Clinical Microbiology
Gain

In two Canadian diagnostic laboratories, AI-based PhenoMATRIX urine culture assessment enabled earlier availability of interpretable results and reduced time to result reporting by about 1.3 hours with automated PM+ release at a tertiary hospital and about 5.3 hours with earlier manual screening at a community lab.

Journal of Clinical Microbiology
Artificial Intelligence for Evidence Synthesis of Emerging Biologics to Improve Skeletal Health in Osteogenesis Imperfecta: Systematic Review and Meta-Analysis
HealthNegative state · G 63 / P 71

GPT-4o-assisted evidence synthesis for osteogenesis imperfecta biologics

Source article: Artificial Intelligence for Evidence Synthesis of Emerging Biologics to Improve Skeletal Health in Osteogenesis Imperfecta: Systematic Review and Meta-Analysis

Problem

GPT-4o showed optimism and positional biases in risk-of-bias assessment because it relied on probabilistic language patterns rather than structured clinical reasoning.

Journal of Medical Internet Research
Gain

GPT-4o integrated into systematic review screening substantially accelerated evidence synthesis for osteogenesis imperfecta biologics while maintaining high sensitivity.

Journal of Medical Internet Research
Preparing students for a world shaped by artificial intelligence | Letters
OtherContested · G 58 / P 61

use of large language models / generative AI by university students affecting deep learning and critical skills

Source article: Preparing students for a world shaped by artificial intelligence | Letters

Problem

Uncritical reliance on generative AI in university coursework risks bypassing deep learning and degrading students' learning in arts and humanities.

The Guardian
Gain

When used thoughtfully in higher education, large language models can enhance teaching and learning by letting students generate and then critique outputs against primary sources.

The Guardian
Beyond EuroSCORE II: is artificial intelligence ready to redefine risk stratification in cardiothoracic surgery?
HealthContested · G 72 / P 68

AI-based risk stratification for cardiothoracic surgery to guide patient selection and perioperative planning

Source article: Beyond EuroSCORE II: is artificial intelligence ready to redefine risk stratification in cardiothoracic surgery?

Problem

AI-based risk stratification tools in cardiothoracic surgery face limited interpretability, dataset bias, inconsistent external validation, and uncertain real-world implementation.

Annals of Medicine & Surgery
Gain

Machine learning models analyzing nonlinear and high-dimensional clinical data have shown improved predictive discrimination for cardiothoracic surgical risk in selected cohorts compared with static traditional scores.

Annals of Medicine & Surgery
Triage safety of patient-facing AI chatbots for nipple discharge: A guideline-informed assessment of red-flag recognition and patient actionability
HealthNegative state · G 65 / P 70

triage safety of AI chatbots answering patient questions about nipple discharge

Source article: Triage safety of patient-facing AI chatbots for nipple discharge: A guideline-informed assessment of red-flag recognition and patient actionability

Problem

A small proportion of chatbot responses were potentially misleading due to missed red-flag features and poor actionability with insufficient action-oriented recommendations.

International Journal of Medical Informatics
Gain

In simulated consultations about nipple discharge, AI chatbots recognized most clinical warning features and were rated safe in most responses.

International Journal of Medical Informatics
Quantifying the impact of slice thickness on cardiovascular risk stratification in lung cancer screening: a multi-center "RESCUE" study
HealthContested · G 68 / P 69

AI-based coronary artery calcium scoring on routine non-gated chest CT for cardiovascular risk stratification

Source article: Quantifying the impact of slice thickness on cardiovascular risk stratification in lung cancer screening: a multi-center "RESCUE" study

Problem

Standard 5.0 mm chest CT reconstructions obscure mild calcification due to partial volume effects, causing significant false-negative CAC zero assessments in routine screening.

Quantitative Imaging in Medicine and Surgery
Gain

Retrospective AI quantification of routinely available thin-slice chest CT reclassifies patients from CAC zero to positive, improving sensitivity for early subclinical atherosclerosis without additional radiation.

Quantitative Imaging in Medicine and Surgery
Performance evaluation of domain-specific and general-purpose AI models for chest radiograph interpretation: a comparative study
HealthPositive state · G 68 / P 63

diagnostic consistency and reporting efficiency for chest radiograph interpretation using M4CXR

Source article: Performance evaluation of domain-specific and general-purpose AI models for chest radiograph interpretation: a comparative study

Problem

Even the domain-specific M4CXR model was inconsistent with reference findings in 25.2% of chest radiograph cases and did not significantly change RADPEER discrepancy rates versus original interpretation.

BMC Medical Imaging
Gain

M4CXR achieved higher report consistency than ChatGPT-4o and cut reporting time from 179.2 seconds unaided to 16.3 seconds assisted when interpreting chest radiographs.

BMC Medical Imaging
Seeing beyond the algorithm: artificial intelligence and the enduring role of the radiologist
HealthContested · G 75 / P 72

AI integration impact on radiology workflow and workload for radiologists

Source article: Seeing beyond the algorithm: artificial intelligence and the enduring role of the radiologist

Problem

AI integration in radiologic practice may paradoxically increase workload and contribute to radiologist burnout when poorly implemented, with automation bias and over-reliance compromising clinical judgment

Current Problems in Diagnostic Radiology
Gain

AI integration in radiologic practice improves triage and reduces report turnaround times while achieving diagnostic performance approaching or exceeding radiologists in chest imaging and breast cancer screening

Current Problems in Diagnostic Radiology
Adopting AI Enhances Humanitarian Operations While Demanding Critical Trade-Offs
HealthContested · G 74 / P 75

adoption of AI tools in humanitarian operations affecting vulnerable communities and humanitarian workers

Source article: Adopting AI Enhances Humanitarian Operations While Demanding Critical Trade-Offs

Problem

Adoption of AI in fragile humanitarian environments creates substantial risk of security breaches from human errors and unregulated data management, and risks reinforcing existing power imbalances for workers and communities.

Avicenna Journal of Medicine
Gain

Adopting AI tools in humanitarian operations was reported to improve health diagnostics, service quality, and analytical efficiency for multifactorial predictions in resource-limited and conflict settings.

Avicenna Journal of Medicine
Custom GPT models for complex rheumatology systematic reviews: A two-part evaluation of data extraction and prognosis appraisal
HealthPositive state · G 73 / P 65

customized GPT-based LLMs for QUIPS risk-of-bias appraisal in rheumatology prognostic systematic reviews

Source article: Custom GPT models for complex rheumatology systematic reviews: A two-part evaluation of data extraction and prognosis appraisal

Problem

GPT-Reviewer showed near-zero agreement with human QUIPS ratings for study participation and outcome measurement, with kappa 0.001.

DIGITAL HEALTH
Gain

Custom GPT models completed all QUIPS domain judgments and reduced data-extraction time from 30.4 to 5.7 minutes per study in rheumatology systematic reviews.

DIGITAL HEALTH