It is 11 p.m. Your four-year-old is finally asleep, and you are on the couch typing a question you were too shy to ask at the last checkup: “Does my kid really need fluoride?” A chatbot answers in seconds — calm, fluent, faintly authoritative. That small, ordinary moment is ChatGPT in pediatric dentistry in a nutshell: instant reassurance, delivered with total confidence. The harder question is whether the confidence is earned.

The 30-Second Version

  • A 2025 study in BMC Oral Health had 30 pediatric dentists grade ChatGPT-4’s answers to 60 common pediatric-dentistry questions.
  • Average accuracy landed around 4 out of 5 — “good” — and the everyday questions parents ask scored just as high as dental-school exam questions.
  • But the model was weakest exactly where stakes are highest: fluoride and baby-tooth nerve (pulp) treatment.
  • The honest caveat: the scores measure how good answers looked to dentists — not whether they are safe to act on without one.

It sounds reassuring. But “sounds right” and “is right” are not the same thing, and pediatric dentistry is full of nuance a confident paragraph can quietly flatten. In a study published in 2025 in BMC Oral Health, Berkant Sezer and Alev Eda Okutan put ChatGPT-4 (the GPT-4o version) to a structured test. They fed it 60 real pediatric-dentistry questions — 30 that parents actually ask and 30 pulled from dental-school teaching and exams — then had 30 pediatric dentists score every answer for accuracy and completeness. The guiding question: can a general-purpose chatbot be trusted with a child’s mouth?


The study, in one glance

The design was deliberately two-sided. Half the questions were the everyday worries clinicians hear from parents — When should my child start brushing? Is a gap between baby teeth normal? Does a root canal on a baby tooth harm the adult tooth underneath? The other half were technical curricular questions spanning six core topics: fissure sealants, fluoride, early childhood caries, oral hygiene, tooth development and occlusion, and pulpal (nerve) therapy. Each answer was rated by 30 pediatric dentists on a 5-point accuracy scale and a 3-point completeness scale, with responses re-scored to steady the ratings.

4.2 / 5
Mean accuracy on parents’ everyday questions
4.45 → 3.93
Best topic (sealants) vs. weakest (baby-tooth pulp therapy)
86.7%
Dentists who’d point patients to AI after seeing the answers
How ChatGPT-4’s accuracy shifts by topicMean expert rating on dental-school questions (5-point scale)at or above 4.0below 4.0‘Good’ threshold = 4.04.54.03.54.454.334.274.003.993.93SealantsEarly cariesHygieneOcclusionFluoridePulp therapyTopic differences on expert questions were statistically significant (p = 0.007).
ChatGPT-4’s mean accuracy on expert (dental-school) questions across six pediatric-dentistry topics: highest on standardized fissure sealants (4.45/5) and lowest on fluoride (3.99) and baby-tooth pulp therapy (3.93). Figure: Decadentry, based on data reported in the study (DOI: 10.1186/s12903-025-06791-9).

Where ChatGPT in pediatric dentistry actually shines

For the bread-and-butter questions, the chatbot was reliably good. Parents’ everyday FAQs averaged 4.21 out of 5 for accuracy; the expert exam questions averaged 4.16 — a difference so small it was not statistically significant (p = 0.942). In plain terms, ChatGPT-4 handled a worried parent’s plain-language question about as well as a dental-school test item. It did best on the most standardized topic of all, fissure sealants (4.45/5, with a median score of a perfect 5). And the professionals were warm to it: after reading the answers, 26 of the 30 dentists (86.7%) said they would be comfortable pointing patients toward AI for information, and 70% would use it in education. For orienting yourself at midnight — what a space maintainer does, when brushing should start — that is a genuine, accessible win. It echoes an earlier Decadentry look at AI in dental education, where the tools beat a lecture but never a human tutor.

So where does it slip?

The comfortable average hides a sharper pattern. Among the expert questions, only two of the six topics fell below the “good” line of 4.0 — fluoride (3.99) and pulpal (nerve) therapy for baby teeth (3.93, the lowest of all). The spread from best topic to worst was statistically significant (p = 0.007), and post-hoc tests confirmed sealants beat fluoride, occlusion, and pulp therapy specifically. The authors’ reading is telling: fluoride is a genuinely contested public topic tangled in conflicting narratives, and baby-tooth pulp therapy demands multifactorial, case-by-case judgment. Those are precisely the situations where a smooth, generic paragraph is most likely to mislead.

⚠ Strong on the simple, shakiest on the high-stakes

The two topics where ChatGPT-4 was weakest — fluoride and pulp therapy — are also two where a wrong or incomplete answer can do real harm. Fluoride has genuine toxicity thresholds that depend on a child’s age and weight; a mismanaged baby-tooth nerve can damage the permanent tooth forming beneath it. Acing the easy 90% is not the same as being safe on the hard 10%.

It behaves like a fluent tour guide who has memorized the main boulevards but starts improvising the moment you reach the back alleys — right when you need the alley.

There is a second catch buried in the method. “Accuracy” and “completeness” here are experts’ subjective impressions of the answers, not verified clinical outcomes. Completeness topped out at 3 (“comprehensive”), yet most answers hovered around 2.5 — covering the essentials but often at the bare minimum, not the fuller picture a clinician would volunteer. The parent questions were also reconstructed by dentists rather than collected from real parents, which the authors concede may have favored simpler, familiar questions and nudged the scores upward. Evaluators were not formally calibrated, inter-rater agreement was not reported, and only one model at one moment in June 2024 was tested. Generative answers drift with version, wording, and time.

Who is accountable when a fluent answer is wrong?

A chatbot holds no license, owes no duty of care, and cannot see the mouth in question. It cannot spot the cavity, weigh the medical history, or take responsibility for what it recommends. The researchers are blunt about it: high Likert scores “should not be taken as indicators of clinical safety or readiness for autonomous use.” The reassuring tone is a property of the technology, not a measure of its correctness — and confidence without accountability is exactly the mix that can talk a tired parent out of calling the dentist.

Could this widen the gap, or narrow it?

The access upside is real. Families without an easy path to a dentist, or who feel judged asking “basic” questions, get a free, patient, non-judgmental explainer at midnight in language they understand. But the same tool can deepen divides. The answers were tested only in English, and the study notes performance can vary by language and phrasing. The families most dependent on a free chatbot are often the least able to afford the specialist visit that would catch its errors. A tool that is “a helpful extra” for the well-resourced and “the only option” for the underserved can quietly create two standards of care.

What this means for you

If you’re a patient (or parent)

Treat it as a starting point, not a verdict. It’s fine for getting your bearings — what a sealant is, when to start brushing. But for anything involving fluoride amounts, a baby-tooth nerve, an injury, or a specific symptom, take the answer as a question to bring to your dentist, not an instruction to follow.

If you’re a clinician

The demand is already here — parents are asking chatbots before and after they ask you. Rather than dismiss it, get ahead of it: point families to vetted resources, ask what they’ve already read, and be ready to correct the confident-but-thin answer, especially on fluoride and pulp therapy.

The bottom line

ChatGPT-4 passed the easy exam and stumbled on the hard one — exactly what you would expect from a tool that pattern-matches language rather than reasons about a specific child. It is a capable explainer and a poor clinician. Used as an assistant that drafts the first paragraph, and never as the oracle that writes the prescription, it can make dentistry more accessible without making it less safe. The instant the question turns hard, the human has to step back in.

Frequently asked questions

Can ChatGPT safely answer questions about my child’s teeth?

For general, well-established topics it did well — roughly 4 out of 5 for accuracy in this study. But it was weakest on fluoride and baby-tooth nerve treatment, and the scores reflect how good answers looked to dentists, not whether they are safe to act on. Use it to get oriented, then confirm anything specific with your dentist.

What did the study actually measure?

Thirty pediatric dentists graded ChatGPT-4’s answers to 60 questions — 30 parent FAQs and 30 dental-school questions — across six topics, rating accuracy out of 5 and completeness out of 3. Parents’ questions averaged 4.21/5 for accuracy; the weakest topic, baby-tooth pulp therapy, scored 3.93/5.

Why was it worse on fluoride and pulp therapy?

The authors suggest fluoride is a publicly debated topic surrounded by conflicting narratives, while baby-tooth pulp therapy requires case-by-case clinical judgment. Both are hard to answer well in a single generic paragraph, which is where a language model tends to be weakest.

Does a high accuracy score mean it’s reliable?

Not on its own. The scores are experts’ subjective ratings, completeness was often only “adequate,” the questions may have skewed simple, and only one model version was tested at one point in time. The authors explicitly warn against treating the results as proof of clinical safety.

Should my dentist be using AI like this?

As a support tool, possibly — for drafting explanations or patient education, with a professional checking the output. Most dentists in the study (about 87%) were open to recommending AI for information. It is an assistant, not a replacement for clinical judgment.

“ChatGPT aced the easy questions about your child’s teeth and wobbled on the hard ones — proof it’s a fluent explainer, not a clinician.”

Source & author credit

This article interprets, and does not reproduce, the following peer-reviewed study. All figures are the authors’ original findings.

Sezer B, Okutan AE. Evaluation of ChatGPT-4’s performance on pediatric dentistry questions: accuracy and completeness analysis. BMC Oral Health. 2025;25(1):1427. DOI: 10.1186/s12903-025-06791-9

ORCID — Berkant Sezer 0000-0001-9731-6156; Alev Eda Okutan 0000-0001-9399-5761

Published open access under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 (CC BY-NC-ND 4.0) license; this article is an independent interpretation, not a reproduction. Reviewed against the primary source per Decadentry’s editorial standards. Decadentry is an independent educational publication and is not affiliated with the study’s authors.

HB

Hossein Boustani Hezarani

Dentist · AI-in-Healthcare researcher · Founder of Decadentry

Hossein writes Decadentry to translate peer-reviewed dental research into clear, honest, jargon-free reading — celebrating what AI can do for dentistry while asking the hard questions the hype skips. Every article is checked against its primary source.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

More posts