TruaceTracing the truth around AIMonday, August 17, 2026
Health·The Trace·Dual reading·Published 2026-08-08

accuracy and sensitivity for orthodontic extraction vs non-extraction treatment planning

Source article: Assessing the Diagnostic Performance of ChatGPT-5.0 versus Machine Learning in Orthodontics: A Comparative Analysis for Extraction Treatment Planning

Abstract: To make accurate orthodontic extraction decisions, various clinical and cephalometric variables must be evaluated. This study aims to evaluate ChatGPT-5.0's performance in distinguishing orthodontic extraction decisions and to compare it with five supervised machine learning (ML) algorithms. Of 550 retrospectively evaluated orthodontic records, 30 were reserved for calibration, leaving 520 for the main analysis. The reference standard was the consensus treatment decision of three expert orthodontists with more t…

TRV-2026-0689Peer-reviewedPermanent record — cite & verify
Trace impact reading

Contested: both sides are scored from claims and sources, not community votes.

P 67The P score combines the specificity and measured human impact of the grounded problem claim with the strength of this Trace’s cited sources.G 70The G score combines the specificity and measured human impact of the grounded gain claim with the strength of this Trace’s cited sources.
Assessing the Diagnostic Performance of ChatGPT-5.0 versus Machine Learning in Orthodontics: A Comparative Analysis for Extraction Treatment Planning

Hospital Universitari Doctor Peset, València 02 by 19Tarrestnom65. CC BY-SA 4.0 · https://creativecommons.org/licenses/by-sa/4.0

The quick read

A comparative study published August 7, 2026 evaluated ChatGPT-5.0 against five supervised machine learning algorithms for orthodontic extraction decisions. Using 520 cases (42.88% extraction, 57.12% non-extraction) and 23 clinical, cephalometric and photographic variables, with expert consensus as reference, ChatGPT-5.0 achieved 75.77% accuracy and 76.68% sensitivity under 5-fold cross-validation, compared to 78.08% accuracy for XGBoost.

The findings matter because extraction planning requires integrating multiple clinical and cephalometric factors and errors affect treatment outcomes. High sensitivity for detecting extraction need suggests large language models could assist decision support comparably to traditional ML, but the lower overall accuracy and specificity versus XGBoost and reliance on retrospective single-center consensus leaves uncertainty about generalizability, prospective performance and clinical workflow integration.

Main points
  • Study analyzed 520 orthodontic cases, 223 extraction (42.88%) and 297 non-extraction (57.12%), with 30 additional cases reserved for calibration from 550 total.
  • Reference standard was consensus treatment decision of three expert orthodontists with more than 5 years of clinical experience using 23 variables including 13 clinical parameters, 7 cephalometric measurements, and photographs.
  • Performance evaluated with 5-fold cross-validation using accuracy, sensitivity, specificity, precision, F1-score and balanced accuracy with 95% confidence intervals, Cochran's Q and McNemar with Holm-Bonferroni correction.
  • XGBoost had highest overall accuracy at 78.08% and highest specificity at 82.15%, while ChatGPT-5.0 had 75.77% accuracy and 76.68% sensitivity.
Gain

ChatGPT-5.0 achieved 75.77% accuracy and the highest sensitivity at 76.68% for orthodontic extraction decisions, performing comparably to XGBoost and significantly better than random forest, SVM, logistic regression and MLP.

Problem

ChatGPT-5.0 did not achieve the highest overall classification accuracy and had lower specificity than XGBoost, indicating it missed the top performance for correctly identifying non-extraction cases.

The rundown

Researchers retrospectively evaluated 550 orthodontic records, reserving 30 for calibration and analyzing 520 main cases against a consensus decision from three orthodontists with over 5 years experience. They used 23 variables spanning clinical parameters, cephalometric measurements and photographs.

ChatGPT-5.0 was tested with 5-fold cross-validation and compared to XGBoost, random forest, SVM, logistic regression and MLP on accuracy, sensitivity, specificity, precision, F1-score and balanced accuracy. Overall differences were significant at p.001, with pairwise tests showing ChatGPT-5.0 significantly above four ML models and similar to XGBoost.

Sources

Reader signal

How should this claim be treated?

The debate