GPT-4.1 predicting publication of AAST surgical abstracts from poster content alone
Source article: Can AI Predict Publication? Multimodal Large Language Models and the Structural Determinants of Surgical Scholarship
Abstract: BackgroundWhether artificial intelligence can identify publishable scientific work is untested. We evaluated whether a multimodal large language model (MLLM) could predict, from poster content alone, which abstracts at the American Association for the Surgery of Trauma (AAST) Annual Meetings reached publication, and characterized the investigator, institutional, and domain level determinants situating model performance.MethodsWe retrospectively analyzed 260 abstracts from the 2021-2022 AAST Annual Meetings. Bibl…
Contested: both sides are scored from claims and sources, not community votes.
The Journal v. 29, no. 19, May 11, 2017 by U.S. Navy. Naval Support Activity (NSA) Bethesda. Public domain
Researchers tested whether GPT-4.1 could predict publication from poster images alone for 260 abstracts presented at the 2021-2022 AAST Annual Meetings. By September 2026 publication date, 142 had published, and the model achieved 58.5% accuracy overall, with higher accuracy in Violence, Societal, and Behavioral and Hemorrhage, Resuscitation, and Vascular Control, but chance-level performance in other domains.
The result matters because it shows AI can read scientific signal legible on the page but cannot read structural advantage that also determines publication, such as multicenter scaffolding and mentorship. Uncertainty remains about generalizability beyond AAST, beyond 2021-2022, and about demographic inference from public data, plus why domain dependence occurs.
- Retrospective analysis of 260 abstracts from 2021-2022 AAST Annual Meetings with bibliographic confirmation of publication.
- 142 (54.6%) reached publication at a mean of 13.4 months, with multicenter origin the only independent predictor ( P = .02).
- Hispanic investigators were underrepresented among first and senior authors and absent from published Hemorrhage, Resuscitation, and Vascular Control subset.
- Model used six-domain rubric averaged over 30 iterations on poster images; performance fell to chance in Critical Care and Outcomes and Systems, Technology, and Process Optimization.
GPT-4.1 scored poster images alone and predicted which AAST abstracts reached publication with 58.5% overall accuracy, rising to 74.2% in Violence, Societal, and Behavioral and 62.2% in Hemorrhage, Resuscitation, and Vascular Control.
Model performance fell to chance in Critical Care and Outcomes and Systems, Technology, and Process Optimization, where publication depended on institutional factors absent from the poster such as multicenter scaffolding, mentorship, and senior author fluency.
The rundown
The study analyzed 260 posters from the 2021-2022 American Association for the Surgery of Trauma meetings, confirming publication via bibliographic search and scoring each poster image with GPT-4.1 on a six-domain rubric averaged over 30 iterations.
Overall 142 abstracts (54.6%) published at mean 13.4 months; multivariable analysis found multicenter origin as the only independent predictor, while Hispanic investigators were underrepresented among first ( P = .03) and senior ( P = .01) authors and absent from the published Hemorrhage, Resuscitation, and Vascular Control subset.
Findings limited to 260 AAST abstracts from 2021-2022, with investigator demographics inferred from public data and model performance dependent on domain and on structural factors not visible on posters.
Sources
- Peer-reviewedThe American Surgeon™2026-09-14
How should this claim be treated?
ace
The debate