Skip to content

AI in Dentistry Glossary

Plain-English definitions of the terms you’ll meet in research on artificial intelligence in dentistry — from the statistics that judge a model to the imaging it reads. 56 terms, cross-checked and kept up to date.

A

Accuracy

The share of all cases a model classifies correctly. It can look high even when a model misses most cases of a rare condition, so dental AI studies should also report sensitivity and specificity.

Artificial intelligence (AI)

Computer systems that perform tasks normally requiring human judgement, such as recognising patterns in images or text. In dentistry, most current AI tools are machine-learning models trained on radiographs, photos, scans or clinical records.

AUC (area under the ROC curve)

A single number from 0.5 (no better than chance) to 1.0 (perfect) summarising how well a model separates cases from non-cases across every possible decision threshold. Useful for comparison, but it does not tell you how the tool behaves at the threshold actually used in clinic.

B

Bias (risk of bias)

Systematic error that can make a study’s results look better or worse than reality — for example, testing a model on images from the same hospital it was trained on. Tools such as QUADAS-2 and PROBAST are used to rate it.

Bitewing radiograph

An intraoral X-ray showing the crowns of upper and lower back teeth together. It is the standard image for detecting decay between teeth, and a common input for caries-detection AI.

Bland–Altman analysis

A method for checking agreement between two measurement techniques by plotting their differences. The ‘limits of agreement’ show the range within which most differences fall — key when AI measurements are compared with a clinician’s.

C

CBCT (cone-beam computed tomography)

A 3D X-ray technique widely used in dentistry for implant planning, endodontics, orthodontics and surgery. AI models are trained to segment teeth, nerves, bone and canals in CBCT volumes.

Cephalometric analysis

Measurement of standard landmarks on a lateral skull radiograph to assess jaw and tooth relationships, mainly in orthodontics. Automatic landmark detection is one of the most mature dental AI applications.

Clinical decision support (CDS)

Software that gives clinicians information or recommendations at the point of care. Dental AI is usually designed as decision support: the clinician remains responsible for the final decision.

Cohen’s kappa

A statistic for agreement between two raters beyond what chance alone would produce. Values near 0 indicate chance-level agreement; values above about 0.6 are usually considered substantial.

Confusion matrix

A table counting a model’s true positives, false positives, true negatives and false negatives. Most diagnostic statistics, such as sensitivity and specificity, are calculated from it.

Convolutional neural network (CNN)

A deep-learning architecture that learns visual features — edges, textures, shapes — directly from images. CNNs underpin most AI tools that read dental radiographs and photographs.

D

Data augmentation

Artificially expanding a training set by flipping, rotating, cropping or adjusting the brightness of images. It helps models generalise but cannot replace genuinely diverse data.

Deep learning

A branch of machine learning that uses neural networks with many layers to learn patterns from large datasets. Most recent dental imaging AI is deep learning.

Dice coefficient

An overlap score between a model’s segmentation and the reference outline, from 0 (no overlap) to 1 (perfect match). Commonly reported for tooth, lesion and nerve segmentation.

Diffusion model

A type of generative AI that creates images by learning to reverse a gradual noising process. In dentistry it has been used to simulate treatment outcomes such as post-orthodontic facial profiles.

E

Explainable AI (XAI)

Methods that show why a model produced an output — for example, a heat map highlighting the image region that drove a caries prediction. Explanations can build trust, but they can also be convincing while being wrong.

External validation

Testing a model on data from a different hospital, device or population than it was trained on. It is the best early signal of whether a tool will work in your clinic, and many dental AI studies still lack it.

F

F1 score

The harmonic mean of precision and recall (sensitivity). It balances missed findings against false alarms in a single number, and is often used for detection tasks.

False positive / false negative

A false positive is a finding flagged when nothing is there (risking overtreatment); a false negative is a real problem the model misses (risking undertreatment). Which error matters more depends on the clinical task.

Federated learning

Training a shared model across several institutions without moving patient data off-site — each site trains locally and only model updates are combined. It is one way to build more diverse dental AI while protecting privacy.

G

Generalisability

How well a model’s performance holds up beyond the data it was developed on — other patients, clinics, scanners and countries. Single-centre studies tell us little about it.

Generative AI

AI that creates new content — text, images or 3D shapes — rather than only classifying existing data. Examples in dentistry include chatbots, crown-design models and treatment simulations.

Ground truth (reference standard)

The answer a model is judged against, such as an expert panel’s diagnosis, histology or a follow-up outcome. A model can only be as trustworthy as its reference standard.

H

Hallucination (AI)

When a generative model, such as a chatbot, produces fluent but false or unsupported information. It is a major risk when large language models answer clinical questions.

I

Intersection over union (IoU)

The overlap between a predicted area and the reference area divided by their combined area. Used to judge segmentation and detection quality; a detection often ‘counts’ only above a threshold such as 0.5.

Intraoral scanner

A handheld device that captures a 3D digital impression of the teeth and gums. Its scans feed AI used for crown design, orthodontic planning and monitoring.

L

Large language model (LLM)

A generative AI model trained on vast amounts of text to predict and produce language, such as ChatGPT, Gemini or Claude. LLMs can explain and summarise, but need verification before clinical use.

M

Machine learning (ML)

Methods that let computers learn patterns from data instead of following hand-written rules. Deep learning is one family of machine-learning methods.

mAP (mean average precision)

A standard score for object-detection models that combines precision and recall across classes and thresholds. Higher is better; it is common in studies using YOLO-style detectors.

Mean absolute error (MAE)

The average size of a model’s errors, ignoring their direction — for example, the average distance in millimetres between predicted and true landmarks.

Meta-analysis

A statistical method that pools results from several studies to estimate an overall effect. Its conclusions are only as reliable as the included studies.

N

Negative predictive value (NPV)

Of all cases a test calls negative, the proportion that are truly negative. It depends on how common the condition is in the population tested.

O

Object detection

A computer-vision task that finds and labels objects in an image with bounding boxes — for example, marking each carious lesion or implant on a radiograph.

Overfitting

When a model learns quirks of its training data so closely that it performs well in development but poorly on new patients. External validation is the main safeguard.

P

Panoramic radiograph (OPG)

A single 2D X-ray showing both jaws, all teeth and surrounding structures. Large panoramic datasets have made it a popular target for dental AI.

Periapical radiograph

An intraoral X-ray showing a whole tooth from crown to root tip and the surrounding bone. It is central to endodontic and periodontal assessment.

Positive predictive value (PPV / precision)

Of all cases a test calls positive, the proportion that are truly positive. A low PPV means many false alarms.

PRISMA

Reporting guidelines for systematic reviews and meta-analyses (Preferred Reporting Items for Systematic reviews and Meta-Analyses). PRISMA-DTA is the version for diagnostic accuracy reviews.

PROBAST

A tool for assessing risk of bias and applicability in studies that develop or validate prediction models, such as AI models that forecast disease progression.

Q

QUADAS-2

A widely used tool for rating risk of bias and applicability in diagnostic accuracy studies, including studies of AI diagnostic tools.

R

Randomised controlled trial (RCT)

A study that randomly assigns participants to groups to compare an intervention against a control. RCTs provide the strongest evidence that an AI tool changes real outcomes, and are still rare in dental AI.

Retrieval-augmented generation (RAG)

A technique that makes a language model look up trusted documents, such as clinical guidelines, before answering, and ground its response in them. It can reduce — not eliminate — hallucinations.

ROC curve

A plot of sensitivity against the false-positive rate (1 − specificity) at every decision threshold. The area under it is the AUC.

S

Segmentation

Labelling every pixel or voxel that belongs to a structure — a tooth, a lesion, the mandibular canal. It produces outlines rather than boxes.

Sensitivity (recall)

Of all patients who truly have a condition, the proportion the test correctly identifies. High sensitivity means few missed cases.

Software as a Medical Device (SaMD)

Software intended for a medical purpose that is not part of a hardware device — including many diagnostic AI tools. Depending on the country, it may need regulatory clearance before clinical use.

Specificity

Of all patients who truly do not have a condition, the proportion the test correctly rules out. High specificity means few false alarms.

Systematic review

A review that uses a pre-defined, reproducible method to find, appraise and summarise all relevant studies on a question. It sits near the top of the evidence hierarchy when the included studies are sound.

T

Teledentistry

Delivering dental assessment, advice or monitoring remotely, often using patient-taken photos or scans. AI is increasingly used to triage these images.

Test set

Data held back from training and used only once, at the end, to estimate how a model performs. If test data leak into training, performance is overstated.

Training set

The data a model learns from. Its size, diversity and labelling quality largely decide how well — and for whom — the model works.

Transformer

A neural-network architecture built around ‘attention’, which weighs how parts of the input relate to each other. It powers large language models and is increasingly used for medical images.

U

U-Net

A neural-network architecture designed for medical image segmentation, with an encoder that compresses the image and a decoder that rebuilds a pixel-level map. It is a workhorse of dental segmentation research.

V

Validation set

Data used during development to tune a model and choose between versions. It is separate from both the training and the final test set.

Y

YOLO (You Only Look Once)

A family of fast object-detection models that locate and classify objects in a single pass through the network. Frequently used for real-time detection on dental radiographs and intraoral photos.

See these ideas in action in our research explainers. Missing a term? Suggest one.