multimodal LLM risk stratification and trajectory classification of oral lichen planus versus expert consensus
Source article: Trajectory-aware risk stratification of oral lichen planus using a multimodal large language model: a longitudinal diagnostic accuracy study
Abstract: To evaluate the performance of a multimodal large language model (LLM) for longitudinal trajectory classification and risk stratification of oral lichen planus (OLP), compared with expert panel consensus. This retrospective diagnostic accuracy study included 300 patients with histopathologically confirmed OLP and at least 24 months of follow-up. Multimodal longitudinal case profiles (serial clinical records, intraoral photographs, and histopathology reports) were independently assessed by (ChatGPT, OpenAI) and a…
Contested: both sides are scored from claims and sources, not community votes.

"Cytomegalovirus infection - Case 301" by Pulmonary Pathology is licensed under CC BY-SA 2.0. To view a copy of this license, visit https://creativecommons.org/licenses/by-sa/2.0/.
Researchers retrospectively tested ChatGPT on 300 histopathologically confirmed oral lichen planus cases with at least 24 months of follow-up, using serial clinical records, intraoral photographs, and histopathology reports. Compared with blinded expert panel consensus, the model achieved 94.7% accuracy for trajectory classification and 78.8% sensitivity with 99.6% specificity for high-risk detection as of the August 2026 publication.
High specificity and trajectory concordance suggest multimodal LLMs could support longitudinal OLP surveillance by integrating clinical information over time, but the 76.3% accuracy for three-level risk stratification and predominance of downward-shift errors indicate risk underestimation remains. Because the data are retrospective and single-cohort, prospective external validation is still required before clinical use.
- Retrospective diagnostic accuracy study of 300 patients with histopathologically confirmed OLP and at least 24 months of follow-up.
- Multimodal longitudinal case profiles included serial clinical records, intraoral photographs, and histopathology reports assessed by ChatGPT and expert panel blinded to results.
- Expert consensus distribution: 156 stable benign (52.0%), 92 inflammatory progression (30.7%), 52 suspicious malignant evolution (17.3%); risk: 73 low, 161 moderate, 66 high.
- Three-level risk stratification accuracy was 76.3% with most errors representing downward shifts (98.6%).
In 300 patients with histopathologically confirmed OLP and at least 24 months follow-up, a multimodal LLM achieved 94.7% trajectory classification accuracy and 99.6% specificity for detecting expert-defined high-risk cases.
The model missed 21.2% of expert-defined high-risk cases with sensitivity of 78.8%, and three-level risk stratification accuracy was only 76.3% with 98.6% of errors being downward shifts that underestimate risk.
The rundown
The study constructed multimodal longitudinal case profiles from serial clinical records, intraoral photographs, and histopathology reports for 300 patients, independently assessed by ChatGPT and an expert panel as reference standard, both blinded to results.
Primary outcome was sensitivity for detecting expert-defined high-risk cases one-versus-rest; secondary outcomes were overall trajectory classification and three-level risk stratification, with expert consensus providing the class distributions and risk levels.
Retrospective single-cohort design without prospective external validation; findings limited to histopathologically confirmed OLP cases with at least 24 months follow-up and may not generalize.
Sources
- Peer-reviewedScientific Reports2026-08-13
How should this claim be treated?
ace
The debate