TruaceTracing the truth around AITuesday, July 21, 2026
TRV-2026-0355Version 1 · Certified

Written 2026-07-20 09:01:09 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-0355
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-07-20T09:01:09.166264Z
status: published
lens: p_space
sector: health
headline: A framework for human evaluation of large language models in healthcare derived from literature review
dek: With generative artificial intelligence (GenAI), particularly large language models (LLMs), continuing to make inroads in healthcare, assessing LLMs with human evaluations is essential to assuring safety and effectiveness. This study reviews existing literature on human evaluation methodologies for LLMs in healthcare across various medical specialties and addresses factors such as evaluation dimensions, sample types and sizes, selection, and recruitment of evaluators, frameworks and metrics, evaluation process,…
gain_title: (none)
problem_title: Human evaluation practices for LLMs in healthcare show gaps in reliability, generalizability, and applicability, undermining assurance of safety and effectiveness.
trace_subject: (none)
gain_reading: (none)
gain_evidence: (none)
problem_reading: Human evaluation practices for LLMs in healthcare show gaps in reliability, generalizability, and applicability, undermining assurance of safety and effectiveness.
problem_evidence: Our literature review of 142 studies shows gaps in reliability, generalizability, and applicability of current human evaluation practices. | assessing LLMs with human evaluations is essential to assuring safety and effectiveness
quick_read: Published September 28, 2024, this peer-reviewed review in npj Digital Medicine examined 142 studies of human evaluation methods for large language models in healthcare. The authors assessed how evaluations are designed and conducted, including dimensions evaluated, sample characteristics, evaluator recruitment, metrics, and analysis, and found systematic gaps in reliability, generalizability, and applicability.

The finding matters because human evaluation is presented as essential to assuring safety and effectiveness as generative AI makes inroads in healthcare. The authors propose the QUEST framework to address those gaps, but the source presents QUEST as a proposal with principles and workflow phases, not as an empirically validated improvement with measured patient or deployment outcomes.
limitation: 
tag: Evidence-backed problem
key_points: Literature review covered 142 studies across various medical specialties examining human evaluation of LLMs in healthcare. | Review examined evaluation dimensions, sample types and sizes, selection and recruitment of evaluators, frameworks and metrics, evaluation process, and statistical analysis type. | Authors propose QUEST framework covering three workflow phases: Planning, Implementation and Adjudication, and Scoring and Review. | QUEST includes five evaluation principles: Quality of Information, Understanding and Reasoning, Expression Style and Persona, Safety and Harm, and Trust and Confidence.
rundown: The authors reviewed 142 studies of human evaluation methodologies for LLMs in healthcare, analyzing dimensions such as sample types and sizes, evaluator selection and recruitment, frameworks and metrics, evaluation process, and statistical analysis type across medical specialties.

Based on identified gaps, they propose QUEST as a comprehensive and practical framework organized into Planning, Implementation and Adjudication, and Scoring and Review, structured around five principles including Quality of Information, Understanding and Reasoning, Expression Style and Persona, Safety and Harm, and Trust and Confidence.
sources:
- peer_reviewed | npj Digital Medicine | https://doi.org/10.1038/s41746-024-01258-7 | 2024-09-28
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
f7f0c46f3be9c4c46df3585da3615c4484e8bda0bd98095935d869a15c766034
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-0355 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.