TruaceTracing the truth around AIWednesday, August 26, 2026
TRV-2026-0061Retracted — remains on the recordPeer-reviewed

Large language models encode clinical knowledge

Abstract Large language models (LLMs) have demonstrated impressive capabilities, but the bar for clinical applications is high. Attempts to assess the clinical knowledge of models typically rely on automated evaluations based on limited benchmarks. Here, to address these limitations, we present MultiMedQA, a benchmark combining six existing medical question answering datasets spanning professional medicine, research and consumer queries and a new dataset of medical questions searched online, HealthSearchQA. We p…

Health · The Trace — both readings · certified 2026-07-12 · v8 · article view · machine-readable

This record was retracted on 2026-07-13Model backfill: source did not support a publishable AI-impact claim. It remains permanently retrievable here, as does every prior version below.
Current reading — gain

Large language models including PaLM 540-billion parameter model and Flan-PaLM encode clinical knowledge and can answer medical questions spanning professional medicine, research and consumer queries as measured on MultiMedQA.

Current reading — problem

LLM answers to medical questions risk failures in factuality, comprehension, reasoning, and introduce possible harm and bias that are not captured by automated evaluations based on limited benchmarks.

What this doesn’t fix

Article text provided is limited to abstract-level description and does not report quantitative performance, error rates, or mitigation results for harm and bias.

Evidence

Cite this record

Truvace Impact Record TRV-2026-0061, v8: “Large language models encode clinical knowledge.” Truvace, 2026-07-13. /record/TRV-2026-0061 (accessed at citation time). sha256 fe5a733a1c19feb5

Calibration history

Every change to this record since certification, in the open.

  1. Retractedv8fe5a733a1c19

    Model backfill: source did not support a publishable AI-impact claim

  2. Revisedv7c94bf0c05062

    Model backfill: grounded claim, summary, sector, and trace validation

  3. Revisedv68a27d1970270

    Reading revised

  4. Sources changedv5440d828c21cc

    Source set updated

  5. Sources changedv486ac56b4b9a1

    Source set updated

  6. Sources changedv361798451a51e

    Source set updated

  7. Sources changedv231d9080080cf

    Source set updated

  8. Certifiedv1d2e299a11abc

    Certified into the record

Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-0061 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.