TruaceTracing the truth around AIMonday, July 20, 2026
Health·The Trace·Model-prefilled trace·Published 2026-07-20

customized GPT-based LLMs for QUIPS risk-of-bias appraisal in rheumatology prognostic systematic reviews

Source article: Custom GPT models for complex rheumatology systematic reviews: A two-part evaluation of data extraction and prognosis appraisal

Background: Systematic reviews are essential for evidence-based practice but remain resource-intensive, particularly during full-text data extraction and structured risk-of-bias appraisal in prognostic research. These challenges are amplified in complex autoimmune diseases such as systemic lupus erythematosus (SLE). Recent advances in large language models (LLMs) have raised interest in their potential; however, rigorous benchmarking against expert reviewers in real-world rheumatology settings is limited. Object…

TRV-2026-0308Peer-reviewedPermanent record — cite & verify
Trace impact reading

Positive state: both sides are scored from claims and sources, not community votes.

P 65The P score combines the specificity and measured human impact of the grounded problem claim with the strength of this Trace’s cited sources.G 73The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.
Custom GPT models for complex rheumatology systematic reviews: A two-part evaluation of data extraction and prognosis appraisal

"MIHANOVICEVA 3 - rheumatology and physiotherapy" by Miroslav Vajdić is licensed under CC BY-SA 2.0. To view a copy of this license, visit https://creativecommons.org/licenses/by-sa/2.0/.

The quick read

Researchers nested a two-part methodological study within two PROSPERO-registered reviews to test customized GPT models on complex rheumatology evidence synthesis. Fifteen SLE metabolomics studies were used to compare human and GPT data extraction, and nineteen rheumatology prognostic studies were reappraised in 2025 with GPT-Reviewer against adjudicated human QUIPS ratings using weighted kappa.

The findings matter because systematic reviews are resource-intensive and automation could accelerate evidence-based practice, but heterogeneous agreement raises concerns about reliability in prognostic research. Uncertainty remains about how to improve table and supplement handling and calibrate models for specific QUIPS domains before human-in-the-loop deployment can be routine.

Main points
  • Two-part study nested within two PROSPERO-registered reviews evaluated custom GPT models for SLE metabolomics data extraction and rheumatology prognosis appraisal.
  • Data extraction comparison used fifteen full-text SLE metabolomics studies processed by humans and GPT with a shared structured template.
  • Prognosis appraisal reappraised nineteen rheumatology prognostic studies in 2025 using customized ChatGPT model GPT-Reviewer versus adjudicated human QUIPS ratings.
  • Agreement quantified with weighted kappa using quadratic weights with 95% confidence intervals.
Gain

Custom GPT models completed all QUIPS domain judgments and reduced data-extraction time from 30.4 to 5.7 minutes per study in rheumatology systematic reviews.

Problem

GPT-Reviewer showed near-zero agreement with human QUIPS ratings for study participation and outcome measurement, with kappa 0.001.

The rundown

In part one, fifteen full-text SLE metabolomics studies were processed by human reviewers and a customized GPT model using a shared, structured template, with concordance and time per study compared.

In part two, nineteen rheumatology prognostic studies with adjudicated human QUIPS domain ratings Low/Moderate/High were reappraised in 2025 by GPT-Reviewer, yielding kappa 0.129 for study attrition, 0.137 for prognostic factor measurement, 0.286 for statistical analysis/reporting, and 0.681 for study confounding.

What this doesn’t fix

Model performance was limited by poor handling of tables/supplements and need for domain-specific calibration before routine use in complex rheumatology synthesis.

Sources

Reader signal

How should this claim be treated?

The debate