The heatmap blooms orange over the lower-left molar. The software has spoken: this is where I found it. It looks like reasoning. It feels like a colleague leaning over your shoulder and pointing. But a coloured blob on a radiograph is not a reason — it is a location. And the gap between those two things is exactly where the credibility of dental AI currently sits.

The 30-Second Version

  • A 2026 systematic review in the International Dental Journal screened 100 records on explainable AI (XAI) in dentistry and found only 19 that qualified.
  • 14 of those 19 carried a high overall risk of bias — small retrospective datasets, weak reference standards, and a near-total absence of external validation.
  • Just one study ever measured whether the explanation actually improved a human being’s understanding.
  • The honest caveat: this is a verdict on the evidence, not on the technology. XAI hasn’t been shown to fail — it has barely been properly tested.

It sounds like the obvious fix. AI’s biggest problem in medicine is the black box: a model hands down a diagnosis with no rationale, and the clinician is left to obey or ignore it. Explainable AI promises to open the box — Grad-CAM heatmaps for images, SHAP and LIME for tabular data — so a dentist can see why. But Sermporn Thaweesapphithak, Thantrira Porntaveetus and colleagues at Chulalongkorn University asked a sharper question than “is XAI spreading?” They asked whether the studies making trust claims are built well enough for those claims to mean anything. The answer is uncomfortable.


The study, in one glance

This was a PRISMA 2020 systematic review (PROSPERO CRD420251182324) searching PubMed, IEEE Xplore, medRxiv and Ovid for dental XAI studies published from 2015 onwards, with searches executed on 30 August 2025. Crucially, the team didn’t just catalogue what exists. They appraised methodological quality with QUADAS-2 for diagnostic accuracy studies and PROBAST for prediction models — the standard tools for asking whether a study’s design can support its conclusions. The 19 included studies spanned cariology, periodontology, oral pathology, orthodontics and forensic dentistry. Convolutional neural networks dominated; Grad-CAM, SHAP and LIME were the explanation methods of choice.

19
of 100 screened records met the inclusion criteria
14/19
carried a high overall risk of bias
1
study tested whether the explanation helped a human understand
How much dental XAI evidence survives scrutinyScreened records, then risk-of-bias appraisal using QUADAS-2 and PROBAST100 RECORDS SCREENED19 included81 excludedTHE 19 INCLUDED STUDIES, BY RISK OF BIAS14 studieshigh overall risk of bias4low to moderate1 study low risk across all domainsTested on actual humans1of 19 studiesmeasured whether the explanation aided understandingThe recurring flawNear-universal absence of external validation:models trained and tested on data from thesame single source.
Evidence quality of explainable-AI research in dentistry, as appraised by QUADAS-2 and PROBAST. Figure: Decadentry, based on data reported in the study (DOI: 10.1016/j.identj.2026.109626).

The promise: the black box really does have a door now

Start with what’s genuinely impressive, because it is. A decade ago, a deep learning model in dentistry was a sealed unit — you got a probability and nothing else. The 19 studies in this review show a field that has learned to open itself up, and some of the work is striking. Motmaen and colleagues trained a ResNet-50 on 26,956 teeth drawn from 1,184 panoramic radiographs to predict tooth extraction decisions, paired it with activation mapping so clinicians could see which regions drove the call, and outperformed dentists with an AUC of 0.901. The ADEPT randomised controlled trial found that dentists using AssistDent detected markedly more enamel-only proximal caries than dentists working unaided.

That second example matters more than its modest size suggests. It is one of the very few instances in this literature where somebody put the AI in front of real clinicians, in a randomised design, and measured what changed. It was also the only study in the entire review rated as low risk of bias across every domain. Rigour like that is clearly achievable. It is just extraordinarily rare.

Does an explanation make a wrong model right?

No — and this is the review’s central and most bracing point. A Grad-CAM heatmap describes what a model attended to. It says nothing about whether the model learned the right thing in the first place. If a network was trained on a few hundred images from one clinic and never validated anywhere else, a beautiful heatmap doesn’t redeem it. It decorates it.

The appraisal found this pattern almost everywhere. Fourteen of nineteen studies were judged high risk of bias, driven by small single-centre retrospective datasets — often scraped from public repositories — and by “ground truth” labels that came from non-expert annotators or imperfect reference standards. The authors describe the lack of external validation as a near-universal limitation: models trained and tested on data from the same source, which inflates apparent performance and quietly forbids any claim to generalise.

⚠ An explanation is not a validation

These are two different guarantees, and the literature routinely conflates them. Validation asks: does this model work on data it has never seen? Explanation asks: what did this model look at? A system can be perfectly explainable and completely unreliable. The review’s finding is that dental XAI has been optimising hard for the second question while mostly skipping the first.

A heatmap on an unvalidated model is a tour guide who has never left the lobby — fluent, confident, and describing every room from a brochure.

There is a subtler problem underneath. The review notes that the field is overwhelmingly built on post hoc explanation: train an opaque model, then bolt on a technique that approximates why it did what it did. That approximation is itself a model, with its own error. An alternative — inherently interpretable models like decision trees or rule-based systems, transparent by construction — remains largely unexplored in dentistry. For high-stakes calls, the authors suggest, the certifiable transparency of a simpler model may beat the uncertain explanation of a complex one.

If the AI explains itself, who is accountable when it’s wrong?

Here is the finding that ought to travel furthest: of nineteen studies, exactly one evaluated whether an explanation changed a human’s understanding. One. The entire clinical justification for XAI — that it lets a dentist verify, understand and confidently act on a recommendation — rests on an assumption almost nobody has tested.

That vacuum has a direction, and it isn’t neutral. Explanations are persuasive by design. A confident visual rationale is exactly the kind of thing that encourages deference, which means a poorly-founded explanation may not merely fail to help — it may actively erode the scepticism that keeps a clinician safe. Until somebody measures whether these tools improve diagnostic confidence, reduce error rates, or shorten time to decision, “explainable” is a description of the software’s output, not a transfer of responsibility. The dentist still owns the diagnosis, with or without the heatmap.

Whose mouths trained these models?

Generalisability is an equity question wearing a technical costume. When the review flags small, single-centre, retrospective datasets and repeatedly raises concerns about representativeness and external validity, what that means in practice is that nobody knows how these systems behave outside the population they were built on. A caries detector validated on one hospital’s patients may quietly underperform on different ages, different dentition stages, different imaging hardware — and because the model still produces a confident heatmap, that failure would be invisible at the chairside.

The irony is sharp. The most compelling case for dental AI is reaching people who don’t have easy access to a specialist — and the studies reviewed here, built on convenience samples from well-resourced centres, are the least equipped to prove it works for them. The authors’ prescription is unambiguous: larger, prospectively collected, diverse datasets, rigorous external validation, and clinically accepted reference standards. That’s not a technical nicety. It’s the precondition for deploying any of this fairly.

What this means for you

If you’re a patient

If a practice tells you their AI “shows why” it flagged something, that’s a reasonable feature — but it isn’t proof the system is accurate. The useful question isn’t can it explain itself? It’s has it been tested on patients like me, somewhere other than where it was built? Your dentist remains the one accountable for the diagnosis, and should be able to explain it in their own words.

If you’re a clinician

Treat explainability as a usability feature, not a quality certificate. Before adopting a tool, ask two questions: was it externally validated on data from a different centre, and has anyone tested whether its explanations actually improve clinician performance? If the vendor’s evidence is a saliency map and an internal accuracy figure, you’re being shown transparency in place of proof.

The bottom line

Explainability was supposed to be the bridge between an algorithm’s competence and a clinician’s trust. What this review documents is a bridge built enthusiastically from one bank only. The techniques are real, the intentions are sound, and roughly three-quarters of the studies making trust claims are not yet solid enough to support them. That isn’t a reason to reject XAI — it’s a reason to stop treating an explanation as though it were evidence. Trust in a clinical tool is earned by proven reliability; transparency is what makes that reliability legible. It’s a complement, not a substitute. An assistant that shows its working is genuinely more useful than one that doesn’t. But showing its working and being right remain two separate achievements, and only one of them has been demonstrated.

Frequently asked questions

What is explainable AI (XAI) in dentistry?

Explainable AI refers to methods that make an AI model’s decisions understandable to people who aren’t AI specialists. In dentistry, the most common techniques are saliency maps such as Grad-CAM, which highlight the regions of a radiograph that most influenced a prediction, and feature-attribution methods such as SHAP and LIME, which show which clinical variables drove a model’s output. The goal is to let a clinician verify and contextualise an AI recommendation rather than accept it blindly.

Does an AI heatmap mean the diagnosis is correct?

No. A heatmap shows where a model looked, not whether the model learned the right thing. This 2026 systematic review found that 14 of 19 dental XAI studies carried a high overall risk of bias, largely because of small single-centre datasets, weak reference standards and an almost universal lack of external validation. An explanation applied to a poorly validated model can create false confidence rather than genuine insight.

Should dentists avoid AI tools that offer explainability?

No — explainability is preferable to an opaque system. The review’s argument is that explainability should not be mistaken for evidence of accuracy. Practitioners are advised to exercise caution because these tools have not been shown to improve diagnostic accuracy or trust in daily practice, and to ask whether a tool has been externally validated on data from a different centre before relying on it clinically.

Has anyone tested whether AI explanations actually help dentists?

Barely. Of the 19 studies included in the review, only one evaluated the impact of explainability on human understanding. The authors identify this as the field’s most important gap and call for human-centred evaluations measuring outcomes such as diagnostic confidence, time to decision, error detection and user satisfaction in controlled studies with dental professionals.

What would make dental XAI trustworthy?

The review sets out four priorities: larger, prospectively collected and demographically diverse datasets; rigorous external validation on data from different centres; clinically accepted reference standards rather than non-expert annotations; and human-centred evaluation that measures whether explanations genuinely improve clinical decision-making. It also urges researchers to consider inherently interpretable models instead of defaulting to post hoc explanations of black boxes.

“An explanation tells you where the AI looked. It does not tell you whether the AI was right — and in dentistry, only one study has ever checked whether the difference registers with a human.”

Related reading on Decadentry: AI vs. endodontist — who reads a dental X-ray better? — a study that shows what happens when an AI is finally graded against physical reality instead of human opinion.

Source & author credit

This article interprets, and does not reproduce, the following peer-reviewed study. All figures are the authors’ original findings.

Thaweesapphithak S, Thongchotchat V, Alinejad-Rokny H, Samaranayake L, Osathanon T, Porntaveetus T. Explainable Artificial Intelligence in Dentistry: A Systematic Review of Its Trust and Translation. International Dental Journal. 2026;76(4):109626. DOI: 10.1016/j.identj.2026.109626

ORCID — S. Thaweesapphithak: 0000-0002-0327-5320 · T. Porntaveetus: 0000-0003-0145-9801

© 2026 The Authors. Published by Elsevier Inc. on behalf of FDI World Dental Federation under a Creative Commons Attribution (CC BY 4.0) licence. The work was supported by Chulalongkorn University, the Health System Research Institute, Thailand Science Research and Innovation, and the Chulalongkorn University Faculty of Dentistry; the authors declared no conflicts of interest. Decadentry is an independent educational publication and is not affiliated with the study’s authors.

HB

Hossein Boustani Hezarani

Dentist · AI-in-Healthcare researcher · Founder of Decadentry

Hossein writes Decadentry to translate peer-reviewed dental research into clear, honest, jargon-free reading — celebrating what AI can do for dentistry while asking the hard questions the hype skips. Every article is checked against its primary source.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

More posts