TruaceTracing the truth around AIFriday, September 11, 2026
Health·The Trace·Dual reading·Published 2026-09-07

AI-based multiclass detection and Gartland classification of pediatric supracondylar fractures on elbow radiographs

Source article: Dual-Approach AI for Pediatric Supracondylar Fractures: Multiclass Radiograph Classification with Explainable AI and Diagnostic Meta-analysis of AI-Based Computational Approaches

Abstract: Rationale and objectives Pediatric supracondylar fractures (SCFs) are the most common elbow injury in children, yet radiographic diagnosis remains challenging due to complex developmental anatomy, with initially missed fracture rates of 17-77%. Prior artificial intelligence (AI) studies have been limited to binary classification frameworks without Gartland subtype differentiation, and no diagnostic test accuracy meta-analysis specific to supracondylar fractures exists. This study aimed to develop the first multi…

TRV-2026-1008Peer-reviewedPermanent record — cite & verify
Trace impact reading

Contested: both sides are scored from claims and sources, not community votes.

P 70The P score combines the specificity and measured human impact of the grounded problem claim with the strength of this Trace’s cited sources.G 70The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.
Dual-Approach AI for Pediatric Supracondylar Fractures: Multiclass Radiograph Classification with Explainable AI and Diagnostic Meta-analysis of AI-Based Computational Approaches

Hospital Universitari Doctor Peset, València 05 by 19Tarrestnom65. CC BY-SA 4.0 · https://creativecommons.org/licenses/by-sa/4.0

The quick read

Researchers developed a YOLOv11 Nano model to detect pediatric supracondylar fractures and classify Gartland subtypes I-III on 1082 elbow radiographs from 2004-2018, testing three patient-level validation schemes and adding bone segmentation and explainable AI visualizations. They also conducted a PRISMA-DTA meta-analysis of four studies totaling 2232 images comparing CNN and radiomics approaches.

The work matters because missed fracture rates of 17-77% are reported for this common childhood elbow injury, and prior AI work was binary without subtype differentiation. While multiclass accuracy exceeded 90% and segmentation boosted performance, the evidence remains early due to only 21 Type I cases, high heterogeneity (I8=92.9%), and lack of prospective multicenter validation.

Main points
  • Retrospective cohort of 1082 pediatric elbow radiographs (2004-2018) with class distribution No-SCF n=576, Type I n=21, Type II n=125, Type III n=360.
  • Three patient-level validation strategies: random split (80:10:10), stratified 5-fold cross-validation, and leave-one-out cross-validation.
  • Bone segmentation integration consistently improved accuracy across all strategies (multiclass +3.9 throughout; binary +6.2 to +10.2% points).
  • PRISMA-DTA meta-analysis of four studies (2232 images) found pooled sensitivity 92.0% and specificity 85.0% with I8=92.9%.
Gain

YOLOv11 Nano achieved multiclass detection and Gartland I-III classification of pediatric supracondylar fractures with ~91-93% accuracy across validation strategies, improving further with bone segmentation.

Problem

Model generalizability for non-displaced Type I fractures is limited by small sample size, and pooled evidence remains preliminary with substantial heterogeneity across studies.

The rundown

The study used YOLOv11 Nano for four-class classification (No-SCF, Gartland Type I-III) on 1082 radiographs, applying data augmentation to mitigate imbalance and NormEnsemble-HiResCAM for explainable AI visualization showing physeal-centered activation for normal/Type I and fragment-boundary foci for Type III.

Meta-analysis of four studies (2232 images) using random-effects modeling reported pooled sensitivity 92.0% (95% CI: 87.0-96.0%) and specificity 85.0% (95% CI: 77.0-92.0%), with YOLOv11 achieving highest specificity 98.0% and diagnostic odds ratio 365.00, but CNN versus radiomics comparison warrants caution.

What this doesn’t fix

Generalizability is constrained by very small Type I sample and limited evidence base with substantial heterogeneity; authors note prospective multicenter validation is required before clinical deployment.

Sources

Reader signal

How should this claim be treated?

The debate