TruaceTracing the truth around AIWednesday, August 26, 2026
TRV-2026-0061Version 2 · Sources changed

Written 2026-07-12 20:50:56 UTC · current record

Reason for this version

Source set updated

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-0061
version: 2
kind: sources_changed
reason: Source set updated
timestamp: 2026-07-12T20:50:56.944130Z
status: published
lens: p_space
sector: health
headline: Large language models encode clinical knowledge
dek: Abstract Large language models (LLMs) have demonstrated impressive capabilities, but the bar for clinical applications is high. Attempts to assess the clinical knowledge of models typically rely on automated evaluations based on limited benchmarks. Here, to address these limitations, we present MultiMedQA, a benchmark combining six existing medical question answering datasets spanning professional medicine, research and consumer queries and a new dataset of medical questions searched online, HealthSearchQA. We propose a human evaluation framework for model answers along multiple axes including factuality, comprehension, reasoning, possible harm and bias. In addition, we evaluate Pathways Language Model 1 (PaLM, a 540-billion parameter LLM) and its instruction-tuned variant, Flan-PaLM 2 on MultiMedQA. Using a combination of prompting strategies, Flan-PaLM achieves state-of-the-art accu…
gain_reading: (none)
problem_reading: Large language models encode clinical knowledge: We propose a human evaluation framework for model answers along multiple axes including factuality, comprehension, reasoning, possible harm and bias.
limitation: Historical evidence reading: the cited study may be limited by its design, population, period, or setting, and later research may report different effects.
tag: Evidence-backed problem
key_points: Abstract Large language models (LLMs) have demonstrated impressive capabilities, but the bar for clinical applications is high. | Attempts to assess the clinical knowledge of models typically rely on automated evaluations based on limited benchmarks. | Here, to address these limitations, we present MultiMedQA, a benchmark combining six existing medical question answering datasets spanning professional medicine, research and consumer queries and a new dataset of medical questions searched online, HealthSearchQA.
rundown: Abstract Large language models (LLMs) have demonstrated impressive capabilities, but the bar for clinical applications is high. Attempts to assess the clinical knowledge of models typically rely on automated evaluations based on limited benchmarks.

Here, to address these limitations, we present MultiMedQA, a benchmark combining six existing medical question answering datasets spanning professional medicine, research and consumer queries and a new dataset of medical questions searched online, HealthSearchQA. We propose a human evaluation framework for model answers along multiple axes including factuality, comprehension, reasoning, possible harm and bias.
sources:
- peer_reviewed | Nature | https://doi.org/10.1038/s41586-023-06291-2 | 2023-07-12
- peer_reviewed | Nature Medicine | https://doi.org/10.1038/s41591-024-03423-7 | 2025-01-08
prev: d2e299a11abce5f7af461d29f16e89d681ba7455ed2c0ad032e2687007be8ae9
sha256
31d9080080cf0ce5647476b9c07e98c5dd4c882461ad8c321e360b1aca250d69
previous
d2e299a11abce5f7af461d29f16e89d681ba7455ed2c0ad032e2687007be8ae9
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-0061 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.