TRV-2026-1274Certified recordPeer-reviewed

Bias and Reliability of AI-Based Peer Review: A Comparative Study of ChatGPT and Claude Evaluating Scientific Abstracts

Background The use of artificial intelligence (AI) models as reviewers of scientific content raises concerns about potential biases related to author identity and about the reproducibility of their evaluations. We assessed whether AI-based reviewers exhibit gender or geographic bias and evaluated the reproducibility of their scoring of scientific abstracts. Methods We randomly selected 10 general internal medicine journals indexed in the Journal Citation Reports (impact factor ≥ 1.5). For each journal, five orig…

Science · The Trace — both readings · certified 2026-10-04 · v1 · article view · machine-readable

Current reading — gain

In a controlled test of 50 abstracts with fictional author identities, ChatGPT and Claude showed no consistent gender or geographic bias and achieved high scoring reproducibility.

Current reading — problem

Observed high agreement may reflect a restricted score range, and the study did not validate LLM scores against human peer review.

What this doesn’t fix

High agreement may be inflated by narrow scoring, and validity compared to human peer review remains untested.

Evidence

Reader signal

How should this claim be treated?

Cite this record

Truvace Impact Record TRV-2026-1274, v1: “Bias and Reliability of AI-Based Peer Review: A Comparative Study of ChatGPT and Claude Evaluating Scientific Abstracts.” Truvace, 2026-10-04. /record/TRV-2026-1274 (accessed at citation time). sha256 bc6be99179c00fc1…

Calibration history

Every change to this record since certification, in the open. None yet — the reading has held since it entered the record.

  1. Certifiedv1bc6be99179c0…

    Certified into the record

Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-1274 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.