A framework for human evaluation of large language models in healthcare derived from literature review
With generative artificial intelligence (GenAI), particularly large language models (LLMs), continuing to make inroads in healthcare, assessing LLMs with human evaluations is essential to assuring safety and effectiveness. This study reviews existing literature on human evaluation methodologies for LLMs in healthcare across various medical specialties and addresses factors such as evaluation dimensions, sample types and sizes, selection, and recruitment of evaluators, frameworks and metrics, evaluation process,…
Human evaluation practices for LLMs in healthcare show gaps in reliability, generalizability, and applicability, undermining assurance of safety and effectiveness.
Evidence
- Peer-reviewednpj Digital Medicine2024-09-28
How should this claim be treated?
Truvace Impact Record TRV-2026-0355, v1: “A framework for human evaluation of large language models in healthcare derived from literature review.” Truvace, 2026-07-20. /record/TRV-2026-0355 (accessed at citation time). sha256 f7f0c46f3be9c4c4…
Calibration history
Every change to this record since certification, in the open. None yet — the reading has held since it entered the record.
Certified into the record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-0355 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace