Evaluation of the Role of Artificial Intelligence in Treatment Planning and Orthodontics: Narrative Review

Orthodontics Treatment Planning Artificial Intelligence Machine Learning Deep Learning Cephalometrics Clear Aligners

Authors

September 8, 2026

Downloads

AI is incorporated in almost all stages of the digital orthodontic workflow, from the automated analysis of diagnostic records, the simulation of treatment outcomes, and remote monitoring of tooth movement. In a very short duration, there has been an explosion of models that are published, which has outpaced critical appraisal of their clinical reliability. The aim of this study is to assess the current role of AI in orthodontic treatment planning through the synthesis of reported diagnostic and predictive performance across the main clinical domains. Furthermore, we will examine the influence of task complexity and study design on performance, and finally identify the methodological, ethical, and regulatory barriers that continue to limit routine clinical integration. A narrative review was conducted utilising PubMed/MEDLINE, Scopus, and Web of Science available for articles published from January 2009 to June 2026. Priority was given to systematic reviews, scoping reviews, meta-analyses, and primary diagnostic –accuracy or predictive-modelling studies published in English. The primary authors provided performance metrics, which were extracted without pooling and are presented accordingly. Seven domain were identified cephalometric landmark identification, malocclusion classification, extraction decision support, skeletal maturation assessment, orthognathic outcome prediction, clear aligner planning, remote monitoring. The reported performance follows a consistent gradient determined by the structure of the task and not by its clinical relevance. The tooth segmentation task achieved 98% performance, the two-dimensional landmark detection task achieved 98.3% performance, the binary crossbite classification task achieved 98.6% performance, and the extraction decision task achieved an area under the curve of 91.2%. Multiclass tasks fall to 71–91%. Fine-grained spatial tasks are lowest, with successful detection rates of 67.5% within 2 mm for posteroanterior landmarks and a pooled three-dimensional landmarking error of 2.44 mm. performance also declines when models are across multiple centres: cervical vertebral maturation staging accuracy fell from 98% in a single-centre cohort to 71.1–78.3% in a six-centre cohort involving 3,600 images. AI is delivers its most robust and reproducible benefits in reduced analysis time and reduce inter-observer variability rather than in unambiguous diagnostic superiority over experienced clinicians. given limited external validation, dataset homogeneity, inconsistent reporting and unresolved ethical and medico-legal question, AI is best positioned as a decision-support adjunct requiring clinician oversight at every stage. priorities for the field are multicentre externally validated datasets and prospective designs with patient-relevant endpoints, routine use of explainable AI and adherence to established reporting frameworks.