DOI RECORD
Comparative Reliability of Gemini and NotebookLM on AAPD Guideline-Based Pediatric Dentistry Multiple-Choice Questions: Accuracy, Consistency, and Error-Pattern Analysis
Abstract
Abstract Aim: To evaluate the reliability characteristics of a standalone Gemini chat workflow and a document-grounded NotebookLM workflow when answering multiple-choice pediatric dentistry questions derived from the 2024 American Academy of Pediatric Dentistry (AAPD) guideline on vital pulp therapies in primary teeth, with particular emphasis on the effects of retrieval-augmented generation (RAG) on evidence-mapping stability, error structure, and potential translational risk. Methodology: This cross-sectional simulation study included 90 single-best-answer pediatric dentistry multiple-choice questions derived from the AAPD guideline. Each item contained four options (A-D) and a predefined gold-standard answer. Gemini and NotebookLM were each tested using three independent accounts; each account answered all 90 questions, yielding 540 responses in total. The standardized prompt was, "You are an experienced physician," and a new conversation was initiated for each query. For NotebookLM, the 2024 AAPD vital pulp therapy guideline PDF was uploaded as a fixed document source. The primary outcome was majority-vote accuracy, and the secondary outcome was inter-account consistency. In addition to Wald and Wilson 95% confidence intervals, Fleiss' kappa, Pearson chi-squared tests, and Fisher's exact tests, an error-structure framework was developed to classify errors as stable or fluctuating and to examine evidence-to-option mapping failures. Results: Gemini achieved a majority-vote accuracy of 92.22% (83/90; Wald 95% CI: 86.69%-97.76%; Wilson 95% CI: 84.81%-96.18%) and an inter-account consistency of 83.33% (75/90; Wilson 95% CI: 74.31%-89.63%; Fleiss' kappa = 0.829). NotebookLM achieved a majority-vote accuracy of 98.89% (89/90; Wald 95% CI: 96.72%-100.00%; Wilson 95% CI: 93.97%-99.80%) and an inter-account consistency of 100.00% (90/90; Wilson 95% CI: 95.91%-100.00%; Fleiss' kappa = 1.000). For accuracy, Pearson's chi-squared test indicated a statistically significant difference (p = 0.030), whereas Fisher's exact test did not reach the 0.05 threshold (p = 0.064), suggesting that the apparent accuracy advantage should be interpreted cautiously. For consistency, NotebookLM was significantly higher in both tests (both p < 0.001). Gemini produced seven majority-vote errors, including one stable error and six fluctuating errors. NotebookLM produced only one error (Q37), which was stable. Errors were concentrated in items requiring precise evidence mapping, including comparisons of caries removal techniques, quantitative outcomes, treatment success rates, Hall technique effects, and diagnostic differentiation. Conclusions: In this pediatric dentistry guideline-based multiple-choice benchmark, the clearest advantage of RAG was not simply an increase in the number of correct answers, but rather improved cross-account stability and reduced random drift. Nevertheless, document grounding did not fully eliminate evidence-to-option mapping failures. Because stable errors are consistent and may appear evidence-based, they may pose more concealed educational and clinical translational risks. Future evaluations of AI systems in dentistry should move beyond accuracy alone toward an integrated reliability framework encompassing correctness, consistency, error architecture, evidence-mapping behavior, and potential risk.
Go to Main Website