TruaceTracing the truth around AITuesday, July 21, 2026
TRV-2026-0308Version 1 · Certified

Written 2026-07-20 08:46:27 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-0308
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-07-20T08:46:27.209874Z
status: published
lens: trace
sector: health
headline: Custom GPT models for complex rheumatology systematic reviews: A two-part evaluation of data extraction and prognosis appraisal
dek: Background: Systematic reviews are essential for evidence-based practice but remain resource-intensive, particularly during full-text data extraction and structured risk-of-bias appraisal in prognostic research. These challenges are amplified in complex autoimmune diseases such as systemic lupus erythematosus (SLE). Recent advances in large language models (LLMs) have raised interest in their potential; however, rigorous benchmarking against expert reviewers in real-world rheumatology settings is limited. Object…
gain_title: Custom GPT models completed all QUIPS domain judgments and reduced data-extraction time from 30.4 to 5.7 minutes per study in rheumatology systematic reviews.
problem_title: GPT-Reviewer showed near-zero agreement with human QUIPS ratings for study participation and outcome measurement, with kappa 0.001.
trace_subject: customized GPT-based LLMs for QUIPS risk-of-bias appraisal in rheumatology prognostic systematic reviews
gain_reading: Custom GPT models completed all QUIPS domain judgments and reduced data-extraction time from 30.4 to 5.7 minutes per study in rheumatology systematic reviews.
gain_evidence: GPT-Reviewer generated complete domain-level QUIPS judgments for all 19 studies | The mean extraction time was shorter for the GPT model than for human reviewers (5.7 vs. 30.4 minutes per study).
problem_reading: GPT-Reviewer showed near-zero agreement with human QUIPS ratings for study participation and outcome measurement, with kappa 0.001.
problem_evidence: heterogeneous concordance versus adjudicated human ratings
quick_read: Researchers nested a two-part methodological study within two PROSPERO-registered reviews to test customized GPT models on complex rheumatology evidence synthesis. Fifteen SLE metabolomics studies were used to compare human and GPT data extraction, and nineteen rheumatology prognostic studies were reappraised in 2025 with GPT-Reviewer against adjudicated human QUIPS ratings using weighted kappa.

The findings matter because systematic reviews are resource-intensive and automation could accelerate evidence-based practice, but heterogeneous agreement raises concerns about reliability in prognostic research. Uncertainty remains about how to improve table and supplement handling and calibrate models for specific QUIPS domains before human-in-the-loop deployment can be routine.
limitation: Model performance was limited by poor handling of tables/supplements and need for domain-specific calibration before routine use in complex rheumatology synthesis.
tag: Model-prefilled trace
key_points: Two-part study nested within two PROSPERO-registered reviews evaluated custom GPT models for SLE metabolomics data extraction and rheumatology prognosis appraisal. | Data extraction comparison used fifteen full-text SLE metabolomics studies processed by humans and GPT with a shared structured template. | Prognosis appraisal reappraised nineteen rheumatology prognostic studies in 2025 using customized ChatGPT model GPT-Reviewer versus adjudicated human QUIPS ratings. | Agreement quantified with weighted kappa using quadratic weights with 95% confidence intervals.
rundown: In part one, fifteen full-text SLE metabolomics studies were processed by human reviewers and a customized GPT model using a shared, structured template, with concordance and time per study compared.

In part two, nineteen rheumatology prognostic studies with adjudicated human QUIPS domain ratings Low/Moderate/High were reappraised in 2025 by GPT-Reviewer, yielding kappa 0.129 for study attrition, 0.137 for prognostic factor measurement, 0.286 for statistical analysis/reporting, and 0.681 for study confounding.
sources:
- peer_reviewed | DIGITAL HEALTH | https://doi.org/10.1177/20552076261467480 | 2026-01-01
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
c9d2ac77e870379e4afb7d5f53f6d962ed1075b1f9b6675eed265e21bc17cc72
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-0308 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.