Somewhere at the tip of a root in your jaw, a lesion the size of a peppercorn may be quietly eating away at bone. It does not hurt. It does not show up when you bite. And on the small X-ray film your dentist just took, it may be almost perfectly invisible — hidden behind a cheekbone, flattened by the geometry of a two-dimensional image. Now imagine an AI that promises to find it for you.

The 30-Second Version

  • Researchers at King’s College London tested a commercial AI platform (Diagnocat) against two experienced endodontists on 339 teeth, using 3D CBCT scans as the gold standard.
  • The AI found 47.9% of confirmed lesions. The human experts found 65.3% — a statistically significant gap (p < .001).
  • On healthy teeth, both were excellent and near-identical: 97.7% vs 95.4% specificity. The AI is good at saying “nothing here.”
  • The uncomfortable part: humans and AI both missed 12.1% of lesions entirely — a limit of flat X-rays themselves, not just the algorithm.

It sounds like the kind of head-to-head that AI usually wins. It didn’t. In a study published in the International Endodontic Journal in May 2025, Marwa Allihaibi, Garrit Koller and Francesco Mannocci ran a retrospective diagnostic accuracy trial that did something most AI-in-dentistry papers skip: instead of grading the algorithm against human opinion, they graded it against physical reality — matched cone beam CT scans that show the lesion in three dimensions. Against that benchmark, the machine came second. The question worth asking is not “did AI lose?” but why the answer flips depending on who is holding the measuring stick.


The study, in one glance

This was a retrospective diagnostic accuracy study reported to STARD 2015 and PRIDASE 2024 standards, drawing on four prospective clinical trials run at Guy’s and St Thomas’ NHS Foundation Trust in London between 2012 and 2022. The team analysed 339 teeth (796 roots) — overwhelmingly molars, all scheduled for a first root canal treatment. Two experienced, calibrated endodontists read the flat periapical films, blinded. A different pair of endodontists read the small-field CBCT scans that served as the reference standard. And Diagnocat analysed the same films fully automatically at its default confidence threshold, with no cropping, enhancement or human help. CBCT confirmed lesions in 121 teeth (35.7%).

47.9%
of confirmed lesions found by the AI (sensitivity, tooth level)
65.3%
found by expert endodontists on the very same films
339
teeth checked against a 3D CBCT reference standard
Where the lesions went121 teeth had an apical lesion confirmed by CBCT, the 3D reference standardExpert endodontistsreading the flat 2D film79 found65.3% sensitivity42 missedAI platformsame films, fully automated58 found47.9% sensitivity63 missedOn healthy teeth, near-identical97.7%vs95.4%specificity — clinicians vs AI (p = .3, not significant)Missed by BOTH41 teeth(12.1%)the limit of 2D imaging itself, not just the algorithmOverall accuracy: clinicians 86.1% vs AI 78.5% (p < .001). AUC 0.81 vs 0.72. Reference standard: cone beam CT.
Detection of apical radiolucencies by expert endodontists versus a commercial AI platform, with CBCT as the diagnostic benchmark. Figure: Decadentry, based on data reported in the study (DOI: 10.1111/iej.14250).

The promise: an algorithm that is very hard to fool by a healthy tooth

Start with what the AI did well, because it is genuinely useful. Diagnocat’s specificity was 95.4% at the tooth level — statistically indistinguishable from the 97.7% posted by two calibrated specialists. When it said a tooth was clear, it was almost always right. Its overall agreement with the human readers was 89%, which tells you the model has learned to see radiographs in a recognisably human way rather than hallucinating pathology.

More interesting still: the AI was not merely a worse copy of the clinicians. When the researchers pooled both readings, total accuracy rose to 87.9% of teeth — and six teeth (1.8%) were correctly called by the algorithm and missed by the humans. Small, but not nothing. That is the shape of a real assistive tool: a second pair of eyes that fails differently than you do, running in seconds, never tired at 4pm on a Friday, and requiring no extra radiation. In a general practice with no endodontist down the hall, a reliable “probably clear” signal has real value.

So why did it find barely half the lesions?

Because sensitivity is where the story turns. The AI detected 58 of 121 confirmed lesions — 47.9%. The clinicians found 79, or 65.3%. At root level the gap held: 39.5% versus 55.9%. Overall accuracy was 78.5% for the AI against 86.1% for the humans, and both differences cleared p < .001. The AI’s area under the ROC curve was 0.72 against the clinicians’ 0.81 — though notably, that particular difference did not reach statistical significance.

The gap widened exactly where you would expect anatomy to punish a flat image. In the upper arch — shallow palatal vault, splayed molar roots, the maxillary sinus floor sitting right on top of the region of interest — clinicians hit 66.0% sensitivity to the AI’s 41.5%. The authors’ explanation is elegant and slightly humbling: experienced clinicians know the film is lying to them. They mentally correct for distortion, superimposition and bad receptor angles built up over thousands of cases. The model has no such theory of its own limitations. It reads the pixels it was given.

⚠ The benchmark problem — this is the finding that should travel

An earlier study of the same platform reported 92.3% sensitivity. This one reports 47.9%. The tool did not get worse; the grader changed. The earlier work used clinician readings of 2D films as ground truth on 60 teeth. This study used CBCT on 339. When your reference standard is human judgement of a flat image, an AI trained on human-annotated flat images will look superb — because it has learned to agree with the humans, including where they are wrong.

Grading a radiograph-reading AI against radiograph-reading humans is like marking an exam with the answer key the student wrote. High agreement is guaranteed. Accuracy is not.

There is a second lever worth knowing about. When the researchers lowered the platform’s detection threshold from 50% to 30% confidence, sensitivity climbed to 70.3% — better than the clinicians. But specificity collapsed from 95.4% to 78.0%, and overall accuracy fell to 75.2%. That is not a bug; it is the immovable trade-off at the heart of every diagnostic test. You can dial an AI to find more disease only by accepting that it will also invent more. In endodontics, invented disease means unnecessary root canals.

If the AI and the expert disagree, who is accountable?

Here is where the study quietly lands a philosophical punch. The authors note that when Diagnocat caught something the clinicians missed, nobody could explain how. A colleague who spots a lesion you overlooked can point at the film and walk you through it; you either accept the reasoning or you don’t. The algorithm produces a green box and a probability score. There is no argument to evaluate — only a claim to obey or ignore.

That turns a clinical disagreement into a coin flip dressed as evidence. And it cuts both ways: a dentist who overrides a correct AI flag, and a dentist who defers to an incorrect one, are both operating without the thing that makes clinical judgement reviewable. Until these systems can show their working, “AI-assisted” is a description of workflow, not a transfer of responsibility. The clinician still owns the diagnosis.

Who does this actually help — and who gets a worse deal?

The equity case for this technology is real, and it is not the one usually pitched. CBCT is the accurate option, but it costs more, delivers more radiation, and simply isn’t available in most practices or most countries. Flat periapical films remain the global default precisely because they are cheap and everywhere. So a tool that squeezes more diagnostic value out of a 2D image is aimed, in principle, at exactly the patients who will never see a CBCT scanner.

But the study’s own numbers complicate the pitch. These films were read by calibrated specialist endodontists — roughly the best-case human comparator. In a busy general practice, the realistic gap between AI and clinician would likely narrow, and the AI’s steady 95% specificity might genuinely add value. The authors are explicit that this remains untested: they call for prospective research including less experienced operators. Until that exists, we are extrapolating. And the deepest limitation is not about AI at all — 41 teeth were missed by the algorithm and both specialists together. No amount of machine learning recovers information that a two-dimensional image never captured.

What this means for you

If you’re a patient

An AI-reviewed X-ray showing “no lesion” is reassuring but not conclusive — in this study that call missed roughly half of confirmed lesions, and experienced dentists missed a third. If you have persistent symptoms and a clear-looking film, asking whether a 3D scan is warranted is a reasonable, evidence-backed question.

If you’re a clinician

Treat a flagged region as a prompt to look again, and an unflagged one as no evidence of absence. Before trusting any vendor’s accuracy figure, ask the only question that matters: what was the reference standard? Numbers benchmarked against human readings of the same modality are measuring agreement, not truth.

The bottom line

This is not a story about AI failing. It is a story about what happens the first time we hold it to an honest standard. Diagnocat performed roughly as well as a careful human at ruling disease out, and meaningfully worse at ruling it in — while inheriting the blind spots of the very experts who trained it. That is not a scandal; it is a specification. An assistant that fails differently from you is worth having beside you. An oracle that fails invisibly, in the same places you do, while sounding certain, is worth rather less. The measure of AI in dentistry will not be how often it agrees with us. It will be how often it is right — and whether we were brave enough to check.

Frequently asked questions

Can AI detect a tooth infection on an X-ray as well as a dentist?

Not yet, according to this study. Using 3D CBCT scans as the gold standard across 339 teeth, a commercial AI platform detected 47.9% of confirmed apical lesions, while experienced endodontists reading the same flat X-rays detected 65.3% — a statistically significant difference. The AI matched the specialists on correctly identifying healthy teeth (95.4% vs 97.7% specificity).

Why do some studies report much higher accuracy for the same AI tool?

Because of what they compare it against. An earlier study of the same platform reported 92.3% sensitivity, but used clinician readings of 2D X-rays as its reference standard on just 60 teeth. This study used CBCT scans on 339 teeth. An AI trained on human-annotated radiographs will closely match human readings — including their errors — so benchmarking against humans measures agreement rather than accuracy.

What is an apical radiolucency and why is it hard to see?

It is a dark area on an X-ray at the tip of a tooth root, usually signalling infection and bone loss from apical periodontitis. Flat 2D X-rays compress three dimensions into one image, so lesions can be hidden by overlapping roots, dense bone, the maxillary sinus or geometric distortion. In this study, 41 teeth (12.1%) had lesions missed by both the AI and the specialists.

Should I ask for a CBCT scan instead of a normal dental X-ray?

Not routinely. CBCT is more accurate but involves higher cost and more radiation, so it is reserved for cases where it will change management. This study reinforces that CBCT remains the radiographic gold standard for detecting apical lesions — but the decision should be made with your dentist based on your symptoms and findings, not by default.

Is AI useful in endodontics at all, then?

Yes, as an adjunct. The platform’s high specificity makes it a reasonable tool for ruling disease out, it agreed with clinicians 89% of the time, and it correctly identified a small number of teeth the specialists missed. Combining both readings raised overall accuracy to 87.9%. The study’s authors conclude it should be viewed as an assistive tool, not a standalone diagnostic system.

“Held to human opinion, the AI scored 92%. Held to physical reality, it scored 48%. Nothing about the algorithm changed — only the honesty of the test.”

Related reading on Decadentry: Can AI turn a smartphone photo into a cavity detector? — a systematic review with a strikingly similar pattern, where AI excelled at obvious decay and faltered on early lesions.

Source & author credit

This article interprets, and does not reproduce, the following peer-reviewed study. All figures are the authors’ original findings.

Allihaibi M, Koller G, Mannocci F. The detection of apical radiolucencies in periapical radiographs: A comparison between an artificial intelligence platform and expert endodontists with CBCT serving as the diagnostic benchmark. International Endodontic Journal. 2025;58(8):1146–1157. DOI: 10.1111/iej.14250

ORCID — M. Allihaibi: 0000-0002-6493-7658 · G. Koller: 0000-0001-7196-1472 · F. Mannocci: 0000-0002-0560-1054

Source article published by Wiley under a Creative Commons Attribution (CC BY 4.0) licence. The study was supported by a Taif University scholarship; the authors declared no conflicts of interest. Decadentry is an independent educational publication and is not affiliated with the study’s authors or with any AI platform named in this article.

HB

Hossein Boustani Hezarani

Dentist · AI-in-Healthcare researcher · Founder of Decadentry

Hossein writes Decadentry to translate peer-reviewed dental research into clear, honest, jargon-free reading — celebrating what AI can do for dentistry while asking the hard questions the hype skips. Every article is checked against its primary source.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

More posts