HealthContested · G 67 / P 64
Source article: Skyer: a novel benchmark for evaluating the effectiveness of large language models in emergency department triage
Problem
Despite higher scores, the evaluated large language models have identified limitations that prohibit them from replacing human experts for triage in overcrowded emergency departments.
Canadian Journal of Emergency MedicineGain
In 55 realistic pediatric scenarios evaluated by the Skyer benchmark, ChatGPT-4.5-preview and Gemini-2.5_05-06 achieved higher triage accuracy and weighted scores than human triage experts, with acceptable consistency across repeats.
Canadian Journal of Emergency MedicineHealthContested · G 70 / P 74
Source article: Performance evaluation of five major large language models in tuberculosis Q&A systems: A multidimensional assessment of readability, quality, and reliability
Problem
The same five LLMs showed significant differences on several readability indices for TB education texts, creating uneven reading difficulty that may undermine patient understanding and adherence.
Frontiers in Public HealthGain
In a 20-question TB Q&A test generating 100 responses, GPT-5 produced the most suitable patient-education texts as measured by C-PEMAT-P among five LLMs.
DIGITAL HEALTHHealthContested · G 75 / P 73
Source article: Mapping artificial intelligence integration in objective structured clinical examinations: A scoping review
Problem
As of the June 2025 search cutoff, evidence did not support claims that AI delivers precision medical education through OSCEs, with weaker performance in relational, situated, and culturally mediated competencies.
Medical TeacherGain
In health professions OSCEs, AI applications improved grading efficiency, feedback speed, and consistency for structured observable tasks, with personalization as the dominant P4 precision education alignment observed by June 2025.
Medical TeacherHealthNegative state · G 68 / P 73
Source article: Combining pathology artificial intelligence and genomic biomarkers to refine long-term postprostatectomy outcome prediction
Problem
MMAI lost independent prognostic value after adjustment for CAPRA clinical risk score and did not outperform established clinical tools for postprostatectomy outcome prediction.
JNCI: Journal of the National Cancer InstituteGain
In 414 prostatectomy cases, the MMAI score derived from digitized pathology images was associated with long-term biochemical recurrence and metastasis, and when combined with genomic CCP score reached the highest discrimination for metastasis.
JNCI: Journal of the National Cancer InstituteOtherPositive state · G 62 / P 55
Source article: Revealed: landmark Scottish AI project has no prospect of meeting renewables promise
Problem
The Lanarkshire AI datacentre complex has no prospect of meeting its promise of 1GW on-site renewable power, faces an acknowledged power provision issue, and will need to connect to the grid instead of operating behind-the-meter.
The GuardianGain
The Lanarkshire AI datacentre complex was presented as delivering jobs and prosperity while being powered entirely from on-site renewables with up to 1GW of new energy infrastructure by 2030.
The GuardianHealthContested · G 73 / P 69
Source article: Improving turnaround times with artificial intelligence in microbiology
Problem
At the low-volume tertiary care hospital, implementing PhenoMATRIX alone without a timely release workflow was associated with increased time to result reporting due to delays between result availability and reporting.
Journal of Clinical MicrobiologyGain
In two Canadian diagnostic laboratories, AI-based PhenoMATRIX urine culture assessment enabled earlier availability of interpretable results and reduced time to result reporting by about 1.3 hours with automated PM+ release at a tertiary hospital and about 5.3 hours with earlier manual screening at a community lab.
Journal of Clinical MicrobiologyHealthNegative state · G 63 / P 71
Source article: Artificial Intelligence for Evidence Synthesis of Emerging Biologics to Improve Skeletal Health in Osteogenesis Imperfecta: Systematic Review and Meta-Analysis
Problem
GPT-4o showed optimism and positional biases in risk-of-bias assessment because it relied on probabilistic language patterns rather than structured clinical reasoning.
Journal of Medical Internet ResearchGain
GPT-4o integrated into systematic review screening substantially accelerated evidence synthesis for osteogenesis imperfecta biologics while maintaining high sensitivity.
Journal of Medical Internet ResearchOtherContested · G 58 / P 61
Source article: Preparing students for a world shaped by artificial intelligence | Letters
Problem
Uncritical reliance on generative AI in university coursework risks bypassing deep learning and degrading students' learning in arts and humanities.
The GuardianGain
When used thoughtfully in higher education, large language models can enhance teaching and learning by letting students generate and then critique outputs against primary sources.
The GuardianHealthContested · G 72 / P 68
Source article: Beyond EuroSCORE II: is artificial intelligence ready to redefine risk stratification in cardiothoracic surgery?
Problem
AI-based risk stratification tools in cardiothoracic surgery face limited interpretability, dataset bias, inconsistent external validation, and uncertain real-world implementation.
Annals of Medicine & SurgeryGain
Machine learning models analyzing nonlinear and high-dimensional clinical data have shown improved predictive discrimination for cardiothoracic surgical risk in selected cohorts compared with static traditional scores.
Annals of Medicine & SurgeryHealthNegative state · G 65 / P 70
Source article: Triage safety of patient-facing AI chatbots for nipple discharge: A guideline-informed assessment of red-flag recognition and patient actionability
Problem
A small proportion of chatbot responses were potentially misleading due to missed red-flag features and poor actionability with insufficient action-oriented recommendations.
International Journal of Medical InformaticsHealthContested · G 68 / P 69
Source article: Quantifying the impact of slice thickness on cardiovascular risk stratification in lung cancer screening: a multi-center "RESCUE" study
Problem
Standard 5.0 mm chest CT reconstructions obscure mild calcification due to partial volume effects, causing significant false-negative CAC zero assessments in routine screening.
Quantitative Imaging in Medicine and SurgeryGain
Retrospective AI quantification of routinely available thin-slice chest CT reclassifies patients from CAC zero to positive, improving sensitivity for early subclinical atherosclerosis without additional radiation.
Quantitative Imaging in Medicine and SurgeryHealthPositive state · G 68 / P 63
Source article: Performance evaluation of domain-specific and general-purpose AI models for chest radiograph interpretation: a comparative study
Problem
Even the domain-specific M4CXR model was inconsistent with reference findings in 25.2% of chest radiograph cases and did not significantly change RADPEER discrepancy rates versus original interpretation.
BMC Medical ImagingGain
M4CXR achieved higher report consistency than ChatGPT-4o and cut reporting time from 179.2 seconds unaided to 16.3 seconds assisted when interpreting chest radiographs.
BMC Medical ImagingHealthContested · G 75 / P 72
Source article: Seeing beyond the algorithm: artificial intelligence and the enduring role of the radiologist
Problem
AI integration in radiologic practice may paradoxically increase workload and contribute to radiologist burnout when poorly implemented, with automation bias and over-reliance compromising clinical judgment
Current Problems in Diagnostic RadiologyGain
AI integration in radiologic practice improves triage and reduces report turnaround times while achieving diagnostic performance approaching or exceeding radiologists in chest imaging and breast cancer screening
Current Problems in Diagnostic RadiologyHealthContested · G 74 / P 75
Source article: Adopting AI Enhances Humanitarian Operations While Demanding Critical Trade-Offs
Problem
Adoption of AI in fragile humanitarian environments creates substantial risk of security breaches from human errors and unregulated data management, and risks reinforcing existing power imbalances for workers and communities.
Avicenna Journal of MedicineGain
Adopting AI tools in humanitarian operations was reported to improve health diagnostics, service quality, and analytical efficiency for multifactorial predictions in resource-limited and conflict settings.
Avicenna Journal of MedicineHealthPositive state · G 73 / P 65
Source article: Custom GPT models for complex rheumatology systematic reviews: A two-part evaluation of data extraction and prognosis appraisal
Problem
GPT-Reviewer showed near-zero agreement with human QUIPS ratings for study participation and outcome measurement, with kappa 0.001.
DIGITAL HEALTHGain
Custom GPT models completed all QUIPS domain judgments and reduced data-extraction time from 30.4 to 5.7 minutes per study in rheumatology systematic reviews.
DIGITAL HEALTH