customized GPT-based LLMs for QUIPS risk-of-bias appraisal in rheumatology prognostic systematic reviews
Source article: Custom GPT models for complex rheumatology systematic reviews: A two-part evaluation of data extraction and prognosis appraisal
Background: Systematic reviews are essential for evidence-based practice but remain resource-intensive, particularly during full-text data extraction and structured risk-of-bias appraisal in prognostic research. These challenges are amplified in complex autoimmune diseases such as systemic lupus erythematosus (SLE). Recent advances in large language models (LLMs) have raised interest in their potential; however, rigorous benchmarking against expert reviewers in real-world rheumatology settings is limited. Object…
Positive state: both sides are scored from claims and sources, not community votes.

"MIHANOVICEVA 3 - rheumatology and physiotherapy" by Miroslav Vajdić is licensed under CC BY-SA 2.0. To view a copy of this license, visit https://creativecommons.org/licenses/by-sa/2.0/.
Researchers nested a two-part methodological study within two PROSPERO-registered reviews to test customized GPT models on complex rheumatology evidence synthesis. Fifteen SLE metabolomics studies were used to compare human and GPT data extraction, and nineteen rheumatology prognostic studies were reappraised in 2025 with GPT-Reviewer against adjudicated human QUIPS ratings using weighted kappa.
The findings matter because systematic reviews are resource-intensive and automation could accelerate evidence-based practice, but heterogeneous agreement raises concerns about reliability in prognostic research. Uncertainty remains about how to improve table and supplement handling and calibrate models for specific QUIPS domains before human-in-the-loop deployment can be routine.
- Two-part study nested within two PROSPERO-registered reviews evaluated custom GPT models for SLE metabolomics data extraction and rheumatology prognosis appraisal.
- Data extraction comparison used fifteen full-text SLE metabolomics studies processed by humans and GPT with a shared structured template.
- Prognosis appraisal reappraised nineteen rheumatology prognostic studies in 2025 using customized ChatGPT model GPT-Reviewer versus adjudicated human QUIPS ratings.
- Agreement quantified with weighted kappa using quadratic weights with 95% confidence intervals.
Custom GPT models completed all QUIPS domain judgments and reduced data-extraction time from 30.4 to 5.7 minutes per study in rheumatology systematic reviews.
GPT-Reviewer showed near-zero agreement with human QUIPS ratings for study participation and outcome measurement, with kappa 0.001.
The rundown
In part one, fifteen full-text SLE metabolomics studies were processed by human reviewers and a customized GPT model using a shared, structured template, with concordance and time per study compared.
In part two, nineteen rheumatology prognostic studies with adjudicated human QUIPS domain ratings Low/Moderate/High were reappraised in 2025 by GPT-Reviewer, yielding kappa 0.129 for study attrition, 0.137 for prognostic factor measurement, 0.286 for statistical analysis/reporting, and 0.681 for study confounding.
Model performance was limited by poor handling of tables/supplements and need for domain-specific calibration before routine use in complex rheumatology synthesis.
Sources
- Peer-reviewedDIGITAL HEALTH2026-01-01
How should this claim be treated?
ace
The debate