Large language models encode clinical knowledge
Abstract Large language models (LLMs) have demonstrated impressive capabilities, but the bar for clinical applications is high. Attempts to assess the clinical knowledge of models typically rely on automated evaluations based on limited benchmarks. Here, to address these limitations, we present MultiMedQA, a benchmark combining six existing medical question answering datasets spanning professional medicine, research and consumer queries and a new dataset of medical questions searched online, HealthSearchQA. We p…
Large language models including PaLM 540-billion parameter model and Flan-PaLM encode clinical knowledge and can answer medical questions spanning professional medicine, research and consumer queries as measured on MultiMedQA.
LLM answers to medical questions risk failures in factuality, comprehension, reasoning, and introduce possible harm and bias that are not captured by automated evaluations based on limited benchmarks.
Article text provided is limited to abstract-level description and does not report quantitative performance, error rates, or mitigation results for harm and bias.
Evidence
- Peer-reviewedNature Medicine2025-01-08
- Peer-reviewedDiagnostics2024-07-09
- Peer-reviewedJournal of Artificial Intelligence, Applications, and Innovations2024-01-01
- Peer-reviewedBioMedInformatics2026-03-13
- Peer-reviewedNature2023-07-12
Truvace Impact Record TRV-2026-0061, v8: “Large language models encode clinical knowledge.” Truvace, 2026-07-13. /record/TRV-2026-0061 (accessed at citation time). sha256 fe5a733a1c19feb5…
Calibration history
Every change to this record since certification, in the open.
Model backfill: source did not support a publishable AI-impact claim
Model backfill: grounded claim, summary, sector, and trace validation
Reading revised
Source set updated
Source set updated
Source set updated
Source set updated
Certified into the record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-0061 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace