TRV-2026-1274Version 1 · Certified

Written 2026-10-04 06:55:10 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-1274
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-10-04T06:55:10.248362Z
status: published
lens: trace
sector: science
headline: Bias and Reliability of AI-Based Peer Review: A Comparative Study of ChatGPT and Claude Evaluating Scientific Abstracts
dek: Background The use of artificial intelligence (AI) models as reviewers of scientific content raises concerns about potential biases related to author identity and about the reproducibility of their evaluations. We assessed whether AI-based reviewers exhibit gender or geographic bias and evaluated the reproducibility of their scoring of scientific abstracts. Methods We randomly selected 10 general internal medicine journals indexed in the Journal Citation Reports (impact factor ≥ 1.5). For each journal, five orig…
gain_title: In a controlled test of 50 abstracts with fictional author identities, ChatGPT and Claude showed no consistent gender or geographic bias and achieved high scoring reproducibility.
problem_title: Observed high agreement may reflect a restricted score range, and the study did not validate LLM scores against human peer review.
trace_subject: LLM scoring of 50 general internal medicine abstracts across fictional gender and geographic identities for quality, novelty, and acceptance
gain_reading: In a controlled test of 50 abstracts with fictional author identities, ChatGPT and Claude showed no consistent gender or geographic bias and achieved high scoring reproducibility.
gain_evidence: no consistent evidence of gender or geographic bias
problem_reading: Observed high agreement may reflect a restricted score range, and the study did not validate LLM scores against human peer review.
problem_evidence: restricted score range may have contributed to the observed agreement | Further studies are needed to assess the validity of LLM-based evaluation and compare it with human peer review
quick_read: In April 2026, researchers tested ChatGPT and Claude as peer reviewers by having each model score 50 general internal medicine abstracts twice under four fictional author identities, producing 800 evaluations of quality, novelty, and acceptance on a 0-10 scale.

By the October 2026 publication date, the study reported no consistent gender or geographic bias and high reproducibility, but noted that narrow score distributions may have inflated agreement and that comparison to human peer review validity was still needed.
limitation: High agreement may be inflated by narrow scoring, and validity compared to human peer review remains untested.
tag: Dual reading
key_points: 50 original research abstracts from 10 general internal medicine journals were each assigned to four fictional identities: African female, African male, American female, American male. | Each abstract was evaluated twice by ChatGPT and Claude on quality, novelty, and acceptance on a 0-10 scale, yielding 800 evaluations in April 2026. | ChatGPT quality and novelty scores were identical across identities with median 7 [1] and 5 [2], and Claude scores were identical across identities for all three dimensions.
rundown: The authors retrieved five original research abstracts from each of 10 general internal medicine journals indexed in Journal Citation Reports with impact factor ≥ 1.5, and created four fictional author identities per abstract. Bias was tested with multivariable ordinal logistic regression, and reproducibility with percent agreement and Fleiss' kappa.

Results showed median quality 7 [1], novelty 5 [2], and acceptance 6-7 [1-2] with minimal variation across identities, and multivariable analyses showed no overall association except for acceptance for Claude at the global level with no individual comparisons statistically significant.
sources:
- peer_reviewed | Journal of General Internal Medicine | https://doi.org/10.1007/s11606-026-10806-8 | 2026-10-02
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
bc6be99179c00fc13fdf7420c91dbb6ac835ea110f3e8742dbf618ecb6a3ca59
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-1274 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.