Accuracy of artificial intelligence in orthodontic extraction treatment planning: a systematic review and meta-analysis
This systematic review and meta-analysis (BMC Oral Health, 2025) evaluated the diagnostic accuracy of AI models in predicting the need for dental extractions during orthodontic treatment planning. Following PRISMA guidelines, seven cross-sectional studies from six countries (India, USA, Chile, China, South Korea, Germany) with 6,261 patients were included. Pooled sensitivity was 70% (95% CI: 61–78) and specificity 90% (95% CI: 87–92). Subgroup analysis showed CNN-based models (ResNet, VGG) had the highest and most consistent performance (sensitivity 76–82%, specificity 93–94%, no heterogeneity), while Random Forest and MLP showed greater variability. Meta-regression indicated higher extraction prevalence was associated with higher sensitivity; funnel plots suggested possible publication bias. AI, especially CNN-based systems, can support orthodontists as a complementary tool for evidence-based extraction decisions but should not replace clinical judgment; large-scale standardized studies are needed for broader implementation.
Background
This systematic review and meta-analysis (BMC Oral Health, 2025) evaluated whether artificial intelligence (AI) can accurately predict the need for dental extractions in orthodontic treatment planning. The protocol was registered on PROSPERO (CRD42024582455). AI and machine learning are increasingly used in dentistry for diagnosis and treatment planning; orthodontic extraction versus non-extraction decisions depend on multiple factors and clinician experience, and AI may help standardise and support such decisions.
Methods
Searches were conducted in PubMed, Scopus, Web of Science, and Google Scholar up to June 2025 using terms related to AI, orthodontics, tooth extraction, and treatment planning. Eligible studies were cross-sectional, assessed AI-based models against clinical standards, and reported diagnostic accuracy (sensitivity and specificity). Data were extracted with a standardised form; quality was assessed with the JBI checklist for analytical cross-sectional studies. Pooled sensitivity and specificity were calculated using a random-effects model; heterogeneity was assessed with I²; subgroup analyses by AI model type and meta-regression (e.g. effect of disease prevalence) were performed; funnel plots were used to assess publication bias.
Results
Seven studies (2021–2024) from six countries were included, with 6,261 patients. Sample sizes ranged from 192 to 1,636. AI techniques included CNNs (ResNet-50/101, VGG16/19), Random Forest, MLP, Decision Trees, SVM, and Auto-WEKA. Pooled sensitivity was 70% (95% CI: 61–78) and pooled specificity 90% (95% CI: 87–92), with high heterogeneity (I² ≈ 97% and 94% respectively). In subgroup analysis, CNN-based models (ResNet and VGG) showed the highest and most consistent performance (e.g. ResNet sensitivity 75.8%, specificity 94.1%; VGG sensitivity 82.4%, specificity 93.1%) with I² = 0%. MLP and Random Forest had moderate performance with substantial heterogeneity. Meta-regression showed that higher extraction prevalence in studies was significantly associated with higher sensitivity (p = 0.050). Funnel plots suggested asymmetry and possible publication bias. Leave-one-out sensitivity analysis indicated that pooled sensitivity was robust.
Limitations and heterogeneity
Heterogeneity was high, likely due to differences in study design, populations, AI algorithms, and datasets. Demographic and clinical detail was often limited. Cross-sectional designs do not allow causal conclusions. The small number of studies and possible publication bias limit the strength of recommendations.
Conclusion and implications
AI models, particularly CNN-based ones, show moderate to high diagnostic accuracy for predicting orthodontic extractions and can be used to support evidence-based decisions. They should serve as complementary tools rather than replacing clinical judgment, especially in complex cases. Standardised training datasets, validation protocols, and large-scale multicentre studies are needed before broader clinical implementation.
Reference: Ziaei S et al. BMC Oral Health. 2025;25:1576. PMC12512631. https://pmc.ncbi.nlm.nih.gov/articles/PMC12512631/
Related reading
Clinical efficacy of combined orthodontic and implant-supported prosthodontic treatment for dentition defects with dentofacial deformities
This retrospective study (Medicine (Baltimore), 2026) evaluated outcomes of combining orthodontic treatment with implant-supported prosthodontic restoration in 104 patients with…
Does Early Orthodontic Treatment in Mixed Dentition Improve Long-Term Outcomes? A Systematic Review and Meta-Analysis
This systematic review and meta-analysis (Medicina, 2025) evaluated whether early orthodontic treatment in children aged 6–12 years during mixed dentition leads to better…
Comparative Study of 0.12% and 0.2% Chlorhexidine Digluconate Mouthwashes on Dental Stains and Gingival Indices
Comparison of the effectiveness of two chlorhexidine concentrations on gum health and its side effects
Recent Advances in Orthodontic Brackets: From Aesthetics to Smart Technologies
This review (Cureus, 2025) explores the latest innovations in orthodontic brackets: aesthetics, function, and smart technology. Ceramic brackets have evolved for better strength…
AI-Driven Advancements in Orthodontics for Precision and Patient Outcomes
This narrative review (Dentistry, MDPI, 2025) explores how artificial intelligence is transforming orthodontic care through personalized treatment planning. AI analyses large…
Assessment of Knowledge and Diagnostic Skills of General Dentists in Khorasan Razavi Province (Iran) Regarding Common Oral Diseases in 2008-2009
Evaluation of knowledge and diagnostic skills of general dentists in the field of common oral mucosal diseases in Khorasan Razavi Province