Diagnostic accuracy of AI-assisted versus independent physician interpretation for bone fractures: a systematic review and meta-analysis
Abstract: Objective We systematically evaluated the diagnostic performance of artificial intelligence (AI)-assisted interpretation versus independent physician assessment for fracture detection. Materials and methods Adhering to PRISMA-DTA guidelines, we searched PubMed and Web of Science for original studies published up to September 17, 2025. Quality was assessed utilizing the QUADAS-3 framework. A bivariate random-effects model pooled diagnostic metrics. Accuracy was assessed by summary receiver operating characteristi…
Hospital Universitari Doctor Peset, València 05 by 19Tarrestnom65. CC BY-SA 4.0 · https://creativecommons.org/licenses/by-sa/4.0
Objective We systematically evaluated the diagnostic performance of artificial intelligence (AI)-assisted interpretation versus independent physician assessment for fracture detection. Materials and methods Adhering to PRISMA-DTA guidelines, we searched PubMed and Web of Science for original studies published up to September 17, 2025.
Compared to unassisted diagnosis, AI significantly improved pooled sensitivity (87%, 95% confidence interval [CI]: 84-89%, versus 73%, 95% CI: 69-78%) and maintained high specificity (95%, 95% CI: 92-97%, versus 94%, 95% CI: 89-96%). Subgroup analysis revealed junior clinicians derived the greatest benefit, exhibiting a 21% absolute sensitivity increase.
- Objective We systematically evaluated the diagnostic performance of artificial intelligence (AI)-assisted interpretation versus independent physician assessment for fracture detection.
- Materials and methods Adhering to PRISMA-DTA guidelines, we searched PubMed and Web of Science for original studies published up to September 17, 2025.
- Quality was assessed utilizing the QUADAS-3 framework.
Compared to unassisted diagnosis, AI significantly improved pooled sensitivity (87%, 95% confidence interval [CI]: 84-89%, versus 73%, 95% CI: 69-78%) and maintained high specificity (95%, 95% CI: 92-97%, versus 94%, 95% CI: 89-96%).
The rundown
Quality was assessed utilizing the QUADAS-3 framework. A bivariate random-effects model pooled diagnostic metrics.
Sources
- Peer-reviewedEuropean Radiology Experimental2026-09-17
How should this claim be treated?
ace
The debate