TruaceTracing the truth around AIMonday, July 20, 2026
TRV-2026-0327Certified recordPeer-reviewed

Skyer: a novel benchmark for evaluating the effectiveness of large language models in emergency department triage

OBJECTIVES: Emergency department (ED) overcrowding causes diagnostic challenges, prolonged wait times, and impairs appropriate triage, often due to human error and fatigue. Large language models can assist ED staff in triage, improving patient care by mitigating these problems. METHODS: We designed an evaluation method (Skyer benchmark) to assess fifteen large language models, including DeepSeek-R1 (70B, 7B), ChatGPT versions (4, 4.5-preview), Gemini iterations (1.5-pro, 2.0-Pro-experimental, 2.5_03-25, 2.5_05-0…

Health · The Trace — both readings · certified 2026-07-20 · v1 · article view · machine-readable

Current reading — gain

In 55 realistic pediatric scenarios evaluated by the Skyer benchmark, ChatGPT-4.5-preview and Gemini-2.5_05-06 achieved higher triage accuracy and weighted scores than human triage experts, with acceptable consistency across repeats.

Current reading — problem

Despite higher scores, the evaluated large language models have identified limitations that prohibit them from replacing human experts for triage in overcrowded emergency departments.

What this doesn’t fix

Authors note limitations that prohibit the best-performing models from replacing human experts, and evaluation was confined to 55 pediatric scenarios with consistency measured over only three repeats.

Evidence

Reader signal

How should this claim be treated?

Cite this record

Truvace Impact Record TRV-2026-0327, v1: “Skyer: a novel benchmark for evaluating the effectiveness of large language models in emergency department triage.” Truvace, 2026-07-20. /record/TRV-2026-0327 (accessed at citation time). sha256 e697c866252df7c1

Calibration history

Every change to this record since certification, in the open. None yet — the reading has held since it entered the record.

  1. Certifiedv1e697c866252d

    Certified into the record

Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-0327 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.