TruaceTracing the truth around AITuesday, July 21, 2026
TRV-2026-0327Version 1 · Certified

Written 2026-07-20 08:46:28 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-0327
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-07-20T08:46:28.035298Z
status: published
lens: trace
sector: health
headline: Skyer: a novel benchmark for evaluating the effectiveness of large language models in emergency department triage
dek: OBJECTIVES: Emergency department (ED) overcrowding causes diagnostic challenges, prolonged wait times, and impairs appropriate triage, often due to human error and fatigue. Large language models can assist ED staff in triage, improving patient care by mitigating these problems. METHODS: We designed an evaluation method (Skyer benchmark) to assess fifteen large language models, including DeepSeek-R1 (70B, 7B), ChatGPT versions (4, 4.5-preview), Gemini iterations (1.5-pro, 2.0-Pro-experimental, 2.5_03-25, 2.5_05-0…
gain_title: In 55 realistic pediatric scenarios evaluated by the Skyer benchmark, ChatGPT-4.5-preview and Gemini-2.5_05-06 achieved higher triage accuracy and weighted scores than human triage experts, with acceptable consistency across repeats.
problem_title: Despite higher scores, the evaluated large language models have identified limitations that prohibit them from replacing human experts for triage in overcrowded emergency departments.
trace_subject: large language model triage performance in 55 pediatric ED scenarios measured by weighted accuracy accounting for over-triage and under-triage
gain_reading: In 55 realistic pediatric scenarios evaluated by the Skyer benchmark, ChatGPT-4.5-preview and Gemini-2.5_05-06 achieved higher triage accuracy and weighted scores than human triage experts, with acceptable consistency across repeats.
gain_evidence: ChatGPT-4.5-preview (77% accuracy, mean weight 377.5 out of 550) and Gemini-2.5_05-06 (74% accuracy, mean weight 365/550) significantly outperformed human-triage-experts accuracy (64% accuracy, mean weight 253.5/550)
problem_reading: Despite higher scores, the evaluated large language models have identified limitations that prohibit them from replacing human experts for triage in overcrowded emergency departments.
problem_evidence: we identified limitations prohibiting these large language models as replacement of human experts
quick_read: Researchers designed the Skyer benchmark to evaluate fifteen large language models on 55 realistic pediatric emergency department scenarios using a weighting system for over-triage and under-triage plus three repeat runs for consistency. By the publication date of July 11 2026, ChatGPT-4.5-preview and Gemini-2.5_05-06 had shown 77% and 74% accuracy with mean weights of 377.5 and 365 out of 550, compared to 64% and 253.5 for human experts.

The finding matters because ED overcrowding impairs triage through human error and fatigue, and a reliable assistive tool could streamline quality care. What remains uncertain is whether performance observed in these 55 pediatric cases generalizes to broader populations and real-world workflows, given the authors' own statement that limitations prohibit replacement of human experts.
limitation: Authors note limitations that prohibit the best-performing models from replacing human experts, and evaluation was confined to 55 pediatric scenarios with consistency measured over only three repeats.
tag: Model-prefilled trace
key_points: Evaluation used 55 realistic clinical pediatric scenarios and a weighting system that accounted for the impacts of over-triage and under-triage | Fifteen models tested including DeepSeek-R1, ChatGPT 4 and 4.5-preview, Gemini 1.5-pro and 2.5 variants, Mistral-7B, Llama-3.3, Gemma, Qwen-2.5, Phi-4-14B | Consistency assessed by repeating tests across scenarios three times; top models showed 85% and 82% consistency
rundown: By July 11 2026, researchers had built Skyer to move beyond simple accuracy by weighting over-triage and under-triage harms and by testing consistency across three repeats of each case. The test set comprised 55 realistic clinical pediatric scenarios applied to fifteen models spanning DeepSeek-R1, ChatGPT, Gemini, Mistral, Llama, Gemma, Qwen and Phi families.

Results reported statistically significant differences with p-value < 0.05 and extremely large effect sizes Cohen's D = 2.18 and 1.98 for the two leaders versus human experts. The authors concluded Skyer selected best-performing models that showed consistent results and potential to assist staff rather than replace them in overcrowded EDs.
sources:
- peer_reviewed | Canadian Journal of Emergency Medicine | https://doi.org/10.1007/s43678-026-01214-2 | 2026-07-11
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
e697c866252df7c1164996087ef56da6c506dfb6efa2b49e405a315f7a9b3895
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-0327 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.