LLM scoring of 50 general internal medicine abstracts across fictional gender and geographic identities for quality, novelty, and acceptance
Source article: Bias and Reliability of AI-Based Peer Review: A Comparative Study of ChatGPT and Claude Evaluating Scientific Abstracts
Background The use of artificial intelligence (AI) models as reviewers of scientific content raises concerns about potential biases related to author identity and about the reproducibility of their evaluations. We assessed whether AI-based reviewers exhibit gender or geographic bias and evaluated the reproducibility of their scoring of scientific abstracts. Methods We randomly selected 10 general internal medicine journals indexed in the Journal Citation Reports (impact factor ≥ 1.5). For each journal, five orig…

Both sides are scored from claims and sources, not community votes.
G 65The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.In brief
In April 2026, researchers tested ChatGPT and Claude as peer reviewers by having each model score 50 general internal medicine abstracts twice under four fictional author identities, producing 800 evaluations of quality, novelty, and acceptance on a 0-10 scale.
By the October 2026 publication date, the study reported no consistent gender or geographic bias and high reproducibility, but noted that narrow score distributions may have inflated agreement and that comparison to human peer review validity was still needed.
Main points
- 50 original research abstracts from 10 general internal medicine journals were each assigned to four fictional identities: African female, African male, American female, American male.
- Each abstract was evaluated twice by ChatGPT and Claude on quality, novelty, and acceptance on a 0-10 scale, yielding 800 evaluations in April 2026.
- ChatGPT quality and novelty scores were identical across identities with median 7 [1] and 5 [2], and Claude scores were identical across identities for all three dimensions.
The gain
In a controlled test of 50 abstracts with fictional author identities, ChatGPT and Claude showed no consistent gender or geographic bias and achieved high scoring reproducibility.
The problem
Observed high agreement may reflect a restricted score range, and the study did not validate LLM scores against human peer review.
The rundown
The authors retrieved five original research abstracts from each of 10 general internal medicine journals indexed in Journal Citation Reports with impact factor ≥ 1.5, and created four fictional author identities per abstract. Bias was tested with multivariable ordinal logistic regression, and reproducibility with percent agreement and Fleiss' kappa.
Results showed median quality 7 [1], novelty 5 [2], and acceptance 6-7 [1-2] with minimal variation across identities, and multivariable analyses showed no overall association except for acceptance for Claude at the global level with no individual comparisons statistically significant.
What this doesn’t fix
High agreement may be inflated by narrow scoring, and validity compared to human peer review remains untested.
Sources
- Peer-reviewedJournal of General Internal Medicine2026-10-02
ace
The debate