You photograph your panoramic X-ray, paste it into ChatGPT, and type: “What’s wrong here?” Ninety seconds later it hands back a confident read — impacted molar, a cyst near the angle of the jaw, the works. A new study on AI dental radiograph reading says an answer like that matches a board-certified radiologist about as often as not. That should thrill you and unsettle you in equal measure.

The 30-Second Version

  • A 2026 Annals of Medicine study pitted three multimodal AI models — ChatGPT, Grok, and a healthcare-tuned model called MANUS — against two board-certified oral radiologists on 120 dental radiographs.
  • On the first pass, ChatGPT and MANUS matched the experts’ diagnosis on 92.5% of images; the radiologists themselves hit 96.7%. A second pass two weeks later nudged the AIs to 90.8–95.0%.
  • The models were strikingly consistent — ChatGPT gave nearly identical reads two weeks apart (κ = 0.937) — and showed no systematic bias against the human benchmark.
  • The honest caveat: every image came from textbooks, not messy real-world scans, and the authors say that likely flatters the AI. No equivalence test was run — “close” was never measured as “as good as.”

It sounds like the radiologist’s job just got automated. But read the methods before you celebrate. The study — by Ahmed Madfa, Abdullah Alshammari and Bassam Anazi at the University of Ha’il, published in Annals of Medicine in 2026 — is careful work, and its own authors are the first to pump the brakes on the headline. The question worth asking isn’t “can AI read a dental radiograph?” It’s “on which radiographs, and compared to what?”


The study, in one glance

The team assembled 120 anonymised radiographs — 40 panoramic (OPG), 40 periapical, and 40 CT slices — drawn from recognised oral-radiology and pathology textbooks, covering tumours, cysts, inflammatory lesions and bony pathology. Two board-certified oral and maxillofacial radiologists, each with 8+ years of experience, set the consensus reference standard. Every image was then fed to ChatGPT (GPT-4-turbo), Grok, and MANUS under identical prompts, and the whole set was re-run two weeks later to test whether each model stayed consistent with itself.

92.5%
ChatGPT & MANUS matched the experts on the first pass (111/120)
96.7%
The human radiologists’ accuracy (116/120)
120
Radiographs tested — all from textbooks, no real patients
How often each reading matched the experts120 curated textbook radiographs · AI second-round scores vs. the human benchmarkRadiologistshuman benchmarkMANUSChatGPTGrok96.7%95.0%93.3%90.8%0255075100%Textbook images are cleaner than real scans — expect lower numbers on everyday clinical radiographs.
Second-round accuracy for the three AI models against the human benchmark, on 120 curated textbook radiographs (radiologists were read once). Figure: Decadentry, based on data reported in the study (DOI: 10.1080/07853890.2026.2664903).

The impressive part: AI dental radiograph reads that rival specialists

Here’s what earns a real double-take. These aren’t bespoke dental AIs trained on millions of X-rays. ChatGPT and Grok are general-purpose chatbots; only MANUS is tuned for healthcare. Yet on this set all three landed within a few points of specialists, and the imaging type — panoramic, periapical, or CT — didn’t significantly change their accuracy. Two weeks later they returned essentially the same answers (ChatGPT’s κ of 0.937 is near-perfect agreement), and a logistic-regression check found no systematic bias — the models weren’t reflexively over-calling or under-calling pathology relative to the humans. For a preliminary triage, a second read, or a teaching aid, that combination of accuracy and consistency is a genuine signal, not noise.

So why shouldn’t you trust it yet?

The catch is in where the pictures came from. Every one of the 120 images was lifted from a textbook — curated, high-resolution, unambiguous, the diagnostic equivalent of flash cards. The authors say it plainly: this “may lead to an overestimation of diagnostic performance compared with real-world clinical settings,” where scans carry artefacts, bad exposures, overlapping anatomy, and patients who inconveniently have more than one thing wrong at once.

⚠ “Close” is not “as good as”

The best model (MANUS, 95.0%) trailed the radiologists (96.7%) by just 1.7 points. Small — but the authors explicitly warn this “should not be interpreted as statistical equivalence,” because no equivalence or non-inferiority test was run. In diagnosis, the gap that matters usually hides in the hard cases, not the average.

A model that aces the textbook is like a student who memorised the practice exam — dazzling until the real questions arrive.

And there’s a deeper worry the study can’t rule out. Textbook radiographs are exactly the kind of images that circulate freely online — and may sit inside a model’s training data. When an AI “recognises” a classic ameloblastoma from a well-known teaching image, is it reading the radiograph, or recalling the caption? The paper doesn’t test for that, and until someone does, headline accuracy on textbook cases deserves an asterisk. As the authors themselves note, these models “rely on probabilistic pattern recognition rather than true clinical reasoning.”

Who’s responsible when the confident answer is wrong?

These models deliver a diagnosis with the same unwavering confidence whether they’re right or hallucinating. On a curated set, that confidence is mostly earned; on your actual scan, it may not be. The authors are unambiguous that final responsibility “must remain with licensed radiologists,” and that regulatory oversight, prospective validation, and medico-legal clarity are prerequisites for real deployment — none of which a chatbot in a browser tab provides. It’s worth remembering what happened when we pitted AI against endodontists on real CBCT scans: accuracy that looks near-human on paper fell sharply once the reference standard got stricter and the images got harder.

Could this widen the gap — or close it?

The optimistic read is access: a clinician in a rural or under-resourced clinic could get a fast second opinion from a tool that costs almost nothing. The pessimistic read is that those are exactly the settings with the messiest images and the least specialist backup to catch a confident AI mistake — the places where an overestimated accuracy number does the most harm. Equity here depends less on the model than on who validates it, and on whose radiographs it was tested against. In this case that was three authors in Saudi Arabia using textbook images; your population, and your scans, may look nothing like the test set.

What this means for you

If you’re a patient

If you paste your X-ray into a chatbot, treat the answer as a prompt for better questions, not a diagnosis. It may be right — it may also be confidently wrong on the one detail that matters. Bring it to your dentist; don’t let it stand in for one.

If you’re a clinician

As a triage aid, a second reader, or a teaching tool, these models are worth watching, and their reproducibility is a real asset. But nothing here validates them on your patients’ real, artefact-laden scans. Keep the licensed human as the final read, and stay skeptical of any accuracy figure built on textbook images.

The bottom line

The headline writes itself: “AI reads X-rays as well as radiologists.” The study underneath is more honest. On a clean textbook set, three AI models came close, stayed consistent, and showed no systematic bias — while their own authors insisted the humans still won, the test was too tidy, and “close” was never measured as “equal.” An assistant that’s fluent in the textbook is a useful colleague. It is not yet the one who should sign off on your care.

Frequently asked questions

Can ChatGPT actually diagnose a dental X-ray?

In this study, ChatGPT matched expert radiologists’ diagnoses on 92.5% of 120 textbook radiographs, rising to 93.3% on a second attempt two weeks later. But those were clean teaching images, not real-world scans, and the authors stress that AI should support — not replace — a qualified clinician.

Was the AI as accurate as a human radiologist?

Close, but not proven equal. The best model, MANUS, scored 95.0% versus 96.7% for the radiologists — a 1.7-point gap. The authors explicitly warn this is not statistical equivalence, because no equivalence or non-inferiority test was performed.

Why does it matter that the images came from textbooks?

Textbook radiographs are curated, high-quality, and unambiguous — much easier than everyday clinical scans full of artefacts and overlapping problems. The authors say this likely overestimates real-world performance, and it raises the untested possibility that the models had already seen these famous images during training.

Which AI model performed best?

MANUS, the healthcare-tuned model, had the highest accuracy (95.0%). ChatGPT was the most consistent between rounds (κ = 0.937). Grok trailed slightly at 90.8%. None showed systematic bias versus the human benchmark, but all three stayed below the radiologists.

Should I paste my own dental X-ray into an AI chatbot?

You can, but treat the result as a conversation starter, not a diagnosis. A chatbot may read your scan correctly or miss the single finding that matters — and it can’t take responsibility for the outcome. Use it to ask your dentist sharper questions, not to skip the visit.

“On the textbook, AI reads like a radiologist. On your actual X-ray, that promise is still unproven — and ‘close’ was never measured as ‘equal.’”

Source & author credit

This article interprets, and does not reproduce, the following peer-reviewed study. All figures are the authors’ original findings. Decadentry’s reporting follows its editorial and fact-checking standards.

Madfa AA, Alshammari AF, Anazi BA. Assessing diagnostic performance of multimodal AI and human experts in oral and maxillofacial radiography: a comparative analysis of ChatGPT, Grok, and MANUS. Annals of Medicine. 2026;58(1):2664903. DOI: 10.1080/07853890.2026.2664903

ORCID — Ahmed A. Madfa: 0000-0001-6124-0129; Abdullah F. Alshammari: 0000-0002-5200-6362

Published open access under a Creative Commons Attribution (CC BY 4.0) licence. Decadentry is an independent educational publication and is not affiliated with the study’s authors.

HB

Hossein Boustani Hezarani

Dentist · AI-in-Healthcare researcher · Founder of Decadentry

Hossein writes Decadentry to translate peer-reviewed dental research into clear, honest, jargon-free reading — celebrating what AI can do for dentistry while asking the hard questions the hype skips. Every article is checked against its primary source.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

More posts