TRV-2026-0355Version 1 · Certified
Reason for this version
Certified into the record
Canonical text (the exact bytes fingerprinted)
TRUVACE RECORD VERSION record: TRV-2026-0355 version: 1 kind: certified reason: Certified into the record timestamp: 2026-07-20T09:01:09.166264Z status: published lens: p_space sector: health headline: A framework for human evaluation of large language models in healthcare derived from literature review dek: With generative artificial intelligence (GenAI), particularly large language models (LLMs), continuing to make inroads in healthcare, assessing LLMs with human evaluations is essential to assuring safety and effectiveness. This study reviews existing literature on human evaluation methodologies for LLMs in healthcare across various medical specialties and addresses factors such as evaluation dimensions, sample types and sizes, selection, and recruitment of evaluators, frameworks and metrics, evaluation process,… gain_title: (none) problem_title: Human evaluation practices for LLMs in healthcare show gaps in reliability, generalizability, and applicability, undermining assurance of safety and effectiveness. trace_subject: (none) gain_reading: (none) gain_evidence: (none) problem_reading: Human evaluation practices for LLMs in healthcare show gaps in reliability, generalizability, and applicability, undermining assurance of safety and effectiveness. problem_evidence: Our literature review of 142 studies shows gaps in reliability, generalizability, and applicability of current human evaluation practices. | assessing LLMs with human evaluations is essential to assuring safety and effectiveness quick_read: Published September 28, 2024, this peer-reviewed review in npj Digital Medicine examined 142 studies of human evaluation methods for large language models in healthcare. The authors assessed how evaluations are designed and conducted, including dimensions evaluated, sample characteristics, evaluator recruitment, metrics, and analysis, and found systematic gaps in reliability, generalizability, and applicability. The finding matters because human evaluation is presented as essential to assuring safety and effectiveness as generative AI makes inroads in healthcare. The authors propose the QUEST framework to address those gaps, but the source presents QUEST as a proposal with principles and workflow phases, not as an empirically validated improvement with measured patient or deployment outcomes. limitation: tag: Evidence-backed problem key_points: Literature review covered 142 studies across various medical specialties examining human evaluation of LLMs in healthcare. | Review examined evaluation dimensions, sample types and sizes, selection and recruitment of evaluators, frameworks and metrics, evaluation process, and statistical analysis type. | Authors propose QUEST framework covering three workflow phases: Planning, Implementation and Adjudication, and Scoring and Review. | QUEST includes five evaluation principles: Quality of Information, Understanding and Reasoning, Expression Style and Persona, Safety and Harm, and Trust and Confidence. rundown: The authors reviewed 142 studies of human evaluation methodologies for LLMs in healthcare, analyzing dimensions such as sample types and sizes, evaluator selection and recruitment, frameworks and metrics, evaluation process, and statistical analysis type across medical specialties. Based on identified gaps, they propose QUEST as a comprehensive and practical framework organized into Planning, Implementation and Adjudication, and Scoring and Review, structured around five principles including Quality of Information, Understanding and Reasoning, Expression Style and Persona, Safety and Harm, and Trust and Confidence. sources: - peer_reviewed | npj Digital Medicine | https://doi.org/10.1038/s41746-024-01258-7 | 2024-09-28 prev: 0000000000000000000000000000000000000000000000000000000000000000
- sha256
- f7f0c46f3be9c4c46df3585da3615c4484e8bda0bd98095935d869a15c766034
- previous
- 0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-0355 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace