TruaceTracing the truth around AIMonday, July 20, 2026
TRV-2026-0355Certified recordPeer-reviewed

A framework for human evaluation of large language models in healthcare derived from literature review

With generative artificial intelligence (GenAI), particularly large language models (LLMs), continuing to make inroads in healthcare, assessing LLMs with human evaluations is essential to assuring safety and effectiveness. This study reviews existing literature on human evaluation methodologies for LLMs in healthcare across various medical specialties and addresses factors such as evaluation dimensions, sample types and sizes, selection, and recruitment of evaluators, frameworks and metrics, evaluation process,…

Health · P Space — documented harm · certified 2026-07-20 · v1 · article view · machine-readable

Current reading — problem

Human evaluation practices for LLMs in healthcare show gaps in reliability, generalizability, and applicability, undermining assurance of safety and effectiveness.

Evidence

Reader signal

How should this claim be treated?

Cite this record

Truvace Impact Record TRV-2026-0355, v1: “A framework for human evaluation of large language models in healthcare derived from literature review.” Truvace, 2026-07-20. /record/TRV-2026-0355 (accessed at citation time). sha256 f7f0c46f3be9c4c4

Calibration history

Every change to this record since certification, in the open. None yet — the reading has held since it entered the record.

  1. Certifiedv1f7f0c46f3be9

    Certified into the record

Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-0355 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.