A framework for human evaluation of large language models in healthcare derived from literature review
With generative artificial intelligence (GenAI), particularly large language models (LLMs), continuing to make inroads in healthcare, assessing LLMs with human evaluations is essential to assuring safety and effectiveness. This study reviews existing literature on human evaluation methodologies for LLMs in healthcare across various medical specialties and addresses factors such as evaluation dimensions, sample types and sizes, selection, and recruitment of evaluators, frameworks and metrics, evaluation process,…
A large aortic aneurism by Shadd, F. J. (Furman Jeremiah), 1852-1908, author. Public domain
Published September 28, 2024, this peer-reviewed review in npj Digital Medicine examined 142 studies of human evaluation methods for large language models in healthcare. The authors assessed how evaluations are designed and conducted, including dimensions evaluated, sample characteristics, evaluator recruitment, metrics, and analysis, and found systematic gaps in reliability, generalizability, and applicability.
The finding matters because human evaluation is presented as essential to assuring safety and effectiveness as generative AI makes inroads in healthcare. The authors propose the QUEST framework to address those gaps, but the source presents QUEST as a proposal with principles and workflow phases, not as an empirically validated improvement with measured patient or deployment outcomes.
- Literature review covered 142 studies across various medical specialties examining human evaluation of LLMs in healthcare.
- Review examined evaluation dimensions, sample types and sizes, selection and recruitment of evaluators, frameworks and metrics, evaluation process, and statistical analysis type.
- Authors propose QUEST framework covering three workflow phases: Planning, Implementation and Adjudication, and Scoring and Review.
- QUEST includes five evaluation principles: Quality of Information, Understanding and Reasoning, Expression Style and Persona, Safety and Harm, and Trust and Confidence.
Human evaluation practices for LLMs in healthcare show gaps in reliability, generalizability, and applicability, undermining assurance of safety and effectiveness.
The rundown
The authors reviewed 142 studies of human evaluation methodologies for LLMs in healthcare, analyzing dimensions such as sample types and sizes, evaluator selection and recruitment, frameworks and metrics, evaluation process, and statistical analysis type across medical specialties.
Based on identified gaps, they propose QUEST as a comprehensive and practical framework organized into Planning, Implementation and Adjudication, and Scoring and Review, structured around five principles including Quality of Information, Understanding and Reasoning, Expression Style and Persona, Safety and Harm, and Trust and Confidence.
Sources
- Peer-reviewednpj Digital Medicine2024-09-28
How should this claim be treated?
ace
The debate