TruaceTracing the truth around AITuesday, August 25, 2026
TRV-2026-0788Version 1 · Certified

Written 2026-08-16 06:22:55 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-0788
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-08-16T06:22:55.518111Z
status: published
lens: trace
sector: health
headline: Trajectory-aware risk stratification of oral lichen planus using a multimodal large language model: a longitudinal diagnostic accuracy study
dek: To evaluate the performance of a multimodal large language model (LLM) for longitudinal trajectory classification and risk stratification of oral lichen planus (OLP), compared with expert panel consensus. This retrospective diagnostic accuracy study included 300 patients with histopathologically confirmed OLP and at least 24 months of follow-up. Multimodal longitudinal case profiles (serial clinical records, intraoral photographs, and histopathology reports) were independently assessed by (ChatGPT, OpenAI) and a…
gain_title: In 300 patients with histopathologically confirmed OLP and at least 24 months follow-up, a multimodal LLM achieved 94.7% trajectory classification accuracy and 99.6% specificity for detecting expert-defined high-risk cases.
problem_title: The model missed 21.2% of expert-defined high-risk cases with sensitivity of 78.8%, and three-level risk stratification accuracy was only 76.3% with 98.6% of errors being downward shifts that underestimate risk.
trace_subject: multimodal LLM risk stratification and trajectory classification of oral lichen planus versus expert consensus
gain_reading: In 300 patients with histopathologically confirmed OLP and at least 24 months follow-up, a multimodal LLM achieved 94.7% trajectory classification accuracy and 99.6% specificity for detecting expert-defined high-risk cases.
gain_evidence: sensitivity was 78.8% (95% CI 67.2-87.5), and specificity was 99.6% (95% CI 97.6-100.0) | Overall trajectory classification accuracy was 94.7% | The multimodal LLM showed high concordance with expert consensus for longitudinal OLP surveillance
problem_reading: The model missed 21.2% of expert-defined high-risk cases with sensitivity of 78.8%, and three-level risk stratification accuracy was only 76.3% with 98.6% of errors being downward shifts that underestimate risk.
problem_evidence: sensitivity was 78.8% (95% CI 67.2-87.5) | most errors representing downward shifts (98.6%) | Three-level risk stratification accuracy was 76.3%
quick_read: Researchers retrospectively tested ChatGPT on 300 histopathologically confirmed oral lichen planus cases with at least 24 months of follow-up, using serial clinical records, intraoral photographs, and histopathology reports. Compared with blinded expert panel consensus, the model achieved 94.7% accuracy for trajectory classification and 78.8% sensitivity with 99.6% specificity for high-risk detection as of the August 2026 publication.

High specificity and trajectory concordance suggest multimodal LLMs could support longitudinal OLP surveillance by integrating clinical information over time, but the 76.3% accuracy for three-level risk stratification and predominance of downward-shift errors indicate risk underestimation remains. Because the data are retrospective and single-cohort, prospective external validation is still required before clinical use.
limitation: Retrospective single-cohort design without prospective external validation; findings limited to histopathologically confirmed OLP cases with at least 24 months follow-up and may not generalize.
tag: Dual reading
key_points: Retrospective diagnostic accuracy study of 300 patients with histopathologically confirmed OLP and at least 24 months of follow-up. | Multimodal longitudinal case profiles included serial clinical records, intraoral photographs, and histopathology reports assessed by ChatGPT and expert panel blinded to results. | Expert consensus distribution: 156 stable benign (52.0%), 92 inflammatory progression (30.7%), 52 suspicious malignant evolution (17.3%); risk: 73 low, 161 moderate, 66 high. | Three-level risk stratification accuracy was 76.3% with most errors representing downward shifts (98.6%).
rundown: The study constructed multimodal longitudinal case profiles from serial clinical records, intraoral photographs, and histopathology reports for 300 patients, independently assessed by ChatGPT and an expert panel as reference standard, both blinded to results.

Primary outcome was sensitivity for detecting expert-defined high-risk cases one-versus-rest; secondary outcomes were overall trajectory classification and three-level risk stratification, with expert consensus providing the class distributions and risk levels.
sources:
- peer_reviewed | Scientific Reports | https://doi.org/10.1038/s41598-026-65667-2 | 2026-08-13
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
148cf95084751790b50ab037bbe1566ac0bbd55cf34eb676ad5b4d268c14afa0
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-0788 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.