Before a surgeon touches a jaw or an orthodontist moves a single tooth, someone has to place the dots. Dozens of tiny anatomical points — the tip of the chin, the base of the nose, the centre of a molar root — pinned onto a skull scan by hand, one careful click at a time. It is slow, it is subjective, and two experts rarely land on exactly the same spot. So what happens when you hand that job to a machine?

The 30-Second Version

  • A multicentre Chinese study trained a lightweight 3D AI to place 41 landmarks on spiral-CT skull scans and 14 on cone-beam CT, hitting an average error of about 1.2 mm across 1,190 scans.
  • The AI matched a senior oral radiologist’s accuracy, comfortably beat a junior one, and cut hands-on marking time by 6 to 9.5 times.
  • It held up on the messy real cases — crooked bites, missing teeth, metal fillings — and almost never “invented” a landmark on a tooth that wasn’t there.
  • The honest caveat: it’s retrospective, built on one region’s scans, and the human-versus-AI face-off used just 40 cases with no blinding — so “clinic-ready” is not yet proven.

It sounds like the drudgery of cephalometric analysis is finally solved. But before we retire the manual tracer, it’s worth reading the fine print. In European Journal of Medical Research, Boyan Liu, Chang Liu, Wei Tang and colleagues at Sichuan University’s West China Hospital of Stomatology built and tested an optimised, lightweight 3D U-Net for AI cephalometric landmark detection — and, refreshingly, they spent as much effort probing where it fails as celebrating where it wins. The guiding question isn’t “can AI place a dot?” It’s “can it place the right dot, on the right patient, reliably enough to plan real surgery?”


The study, in one glance

This was a multicentre retrospective diagnostic study. The model was trained and tested on 480 spiral-CT (SCT) and 240 cone-beam CT (CBCT) cases from West China Hospital, then run on a separate external set of 320 SCT and 150 CBCT scans from two other centres — 1,190 scans in all, from patients aged roughly 6 to 70. Accuracy was measured as mean radial error (MRE): the straight-line distance, in millimetres, between the AI’s dot and the expert’s reference dot. Lower is better, and anything within 2 mm is generally treated as clinically acceptable.

1.22 mm
Average 3D landmark error across all 1,190 scans
≈ Senior
AI matched a senior specialist (1.25 vs 1.33 mm); beat the junior (1.90 mm)
6–9.5×
Faster point-marking with AI assistance
How close to the true point? (mean error, mm)Lower is better — every bar sits inside the clinically acceptable 2 mm limit.2 mm limitJunior1.90Senior1.33AI model1.25Error vs experts on 40 external cases. AI matched the senior specialist and beat the junior.
Mean radial error, AI versus human observers. Figure: Decadentry, based on data reported in the study (DOI: 10.1186/s40001-025-03198-8).

Millimetre precision that held up in the messy cases

The headline result is genuinely impressive, and not just because the numbers are small. Averaged over the whole 1,190-scan dataset the error was 1.22 mm, dropping to about 1.01 mm on the higher-resolution cone-beam scans. More importantly, the model didn’t collapse when the data got ugly. The authors deliberately kept in cases with malocclusion, missing teeth and metal artifacts — exactly the patients a curated demo would quietly exclude — and accuracy stayed below 1.4 mm with no statistically significant drop. Performance also barely changed between the hospital’s own scans and the two external centres, which is the kind of external check that a lot of dental-AI papers skip entirely. And when a tooth genuinely wasn’t there, the model correctly recognised its absence 97.7% of the time rather than hallucinating a landmark into empty gum — a small detail with real safety weight.

But is 1.2 millimetres actually good enough?

Here’s the uncomfortable part: a millimetre isn’t a millimetre everywhere. Cephalometric planning doesn’t use the dots directly — it draws lines and angles between them, and small positional slips get amplified into the measurements that actually drive treatment. The study found the errors weren’t evenly spread. The coronal (depth) axis carried the biggest deviations, tied to CT slice thickness, and specific points behaved worse than the average suggested: the lower incisor drifted up to about 1.7 mm, and the mandibular angle landmark on cone-beam CT reached nearly 1.5 mm. An average of 1.2 mm is reassuring; a 1.7 mm wobble on the exact tooth you’re planning to move is less so.

⚠ The average hides the outliers

Whether sub-2 mm precision is “safe” depends entirely on the procedure. A millimetre of slack barely matters for a broad growth assessment; it can matter a great deal at the incisor edge or the surgical midline. The authors themselves flag that they haven’t yet measured how these landmark errors propagate into the linear and angular values clinicians rely on.

It’s a superb GPS for the skull: it drops the pin to within a millimetre. But you still decide whether that’s the right corner to operate on.

There’s also the question of what the AI was compared against. The reference “truth” was itself drawn by humans — a 9-year and a 14-year expert, quality-checked by a 31-year veteran. That’s a strong standard, but it means the model learned to imitate a particular group’s habits, not some objective anatomical ground truth. An AI that agrees with your experts is valuable; it is not the same as an AI that is right.

Who’s accountable when the pin lands wrong?

The AI-versus-human showdown — the part everyone quotes — rested on just 40 external cases, with no blinding and a known risk of “halo effect,” where evaluators trust a marking more simply because a computer suggested it. The authors are commendably blunt about this: they write that the results “cannot fully substantiate clinical reliability” and call for prospective, randomised validation before anyone adopts it. That candour matters, because the model’s best trick — speeding a junior clinician’s work by nearly 29% — is also its subtlest risk. A tool that quietly nudges a less-experienced eye toward accepting its dots can raise average quality and erode independent judgement at the same time. We saw a version of this tension when AI went head-to-head with endodontists reading X-rays: the software can be fast and confident and still miss what an expert wouldn’t.

Would it work on your face?

Every scan in this study came from three centres in one region of China, with a mean patient age around 30 and a specific set of CT and CBCT machines. Craniofacial anatomy varies with ancestry, age and growth stage; scanner brand and slice thickness change the images themselves. A model this accurate on its home turf could stumble on a 9-year-old’s growing face, an elderly edentulous jaw, or images from a different manufacturer — and none of that has been tested yet. High performance on a narrow population is a promising start, not a passport to every clinic.

What this means for you

If you’re a patient

If your orthodontist or surgeon uses AI-assisted scan analysis, it’s a speed-and-consistency aid, not a replacement for their judgement. The clinician still places, checks and signs off on every point that shapes your plan. It’s entirely reasonable to ask whether AI was involved and who verified the result.

If you’re a clinician

Treat tools like this as a fast first draft that boosts a junior’s baseline and trims tedious marking time — then review every landmark yourself, especially depth-axis and single-tooth points where errors concentrate. Until prospective, multi-population validation exists, keep your hand on the cursor.

The bottom line

This is one of the more honest AI-in-dentistry papers you’ll read: it delivers expert-level precision on hard cases and then spends its discussion listing every reason not to trust it yet. That’s the right posture. An AI that pins skull landmarks to a millimetre is a remarkable assistant — but it’s an assistant, not an oracle. It hands the clinician a faster, steadier starting point; it doesn’t hand over the responsibility for where the lines are finally drawn.

Frequently asked questions

What is cephalometric landmark detection, and why does it matter?

It’s the process of marking specific anatomical points — like the tip of the chin or the centre of a tooth root — on a skull X-ray or CT scan. Orthodontists and jaw surgeons draw lines and angles between these points to plan braces, growth guidance and surgery, so where the points sit directly shapes the treatment.

Is AI now more accurate than a dentist at placing these landmarks?

Not more accurate than an expert. In this study the AI matched a senior specialist (about 1.25 vs 1.33 mm error) and beat a less-experienced junior, but the authors are clear its accuracy does not surpass that of experienced clinicians — and the comparison used only 40 cases without blinding.

Is 1 to 2 millimetres of error safe for planning surgery or braces?

It depends on the procedure. Under 2 mm is generally considered clinically acceptable for many analyses, but the same slack matters more at the incisor edge or a surgical midline than in a broad growth assessment. The study also found certain points and the depth axis carried larger-than-average errors.

Does the AI ever mark landmarks that aren’t actually there?

Rarely. When a tooth was genuinely missing, the model correctly recognised its absence about 97.7% of the time instead of inventing a landmark — an important safety feature, since a hallucinated point could distort the whole plan.

Can I get this AI at my orthodontist today?

Probably not as tested here. It’s a research model validated retrospectively on one region’s scans. The authors explicitly call for prospective, randomised trials across more diverse patients and scanners before clinical adoption.

“AI can drop a pin on your skull to within a millimetre — but a millimetre in the wrong direction is still the surgeon’s call, not the software’s.”

Source & author credit

This article interprets, and does not reproduce, the following peer-reviewed study. All figures are the authors’ original findings.

Liu B, Liu C, Xiong Y, Zhu H, Zeng W, Chen J, Guo J, Liu W, Tang W. Accuracy and reliability of 3D cephalometric landmark detection with deep learning. European Journal of Medical Research. 2025;30(1):1000. DOI: 10.1186/s40001-025-03198-8

ORCID — no author ORCID iDs are listed in the article’s public metadata, so none are cited here.

Published open access under a CC BY 4.0 licence. Interpreted for Decadentry by Hossein Boustani Hezarani under our editorial & fact-checking standards. Decadentry is an independent educational publication and is not affiliated with the study’s authors.

HB

Hossein Boustani Hezarani

Dentist · AI-in-Healthcare researcher · Founder of Decadentry

Hossein writes Decadentry to translate peer-reviewed dental research into clear, honest, jargon-free reading — celebrating what AI can do for dentistry while asking the hard questions the hype skips. Every article is checked against its primary source.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

More posts