Artificial Hallucination in Generative AI: A Comparative Study on Accuracy, Reasoning, and Evidence Reliability in Multiple-Choice Questions



Author Information

Tan Cheng Keat, Nanyang Polytechnic, Singapore
Yin Ni Annie Ng, Nanyang Polytechnic, Singapore
Ng Qing Hao, Lucence, Singapore
Seh Yi Joseph Tan, GlaxoSmithKline Asia House, Singapore

Abstract

The rapid advancement of Large Language Models (LLMs) has spurred the rise of artificial intelligence-generated content (AIGC), heightening concerns about artificial hallucinations - fictional, erroneous, or unsubstantiated information - can foster misconceptions and hinder critical thinking among learners. This is particularly concerning in evidence-based disciplines like pharmacology, where factual precision underpins analysis and application. Ensuring the reliability of AIGC is therefore essential to understanding the educational role of LLM-assisted learning. This study evaluated the accuracy, rationale validity, and citation reliability of four LLMs—ChatGPT-4o, Google Gemini, Microsoft Copilot, and Claude—when answering fifty expert-generated pharmacology multiple-choice questions (MCQs) independently validated by two pharmacists. Using a standardized prompt with sessions reset between questions, each model’s responses were assessed for accuracy, reasoning validity, and citation relevance. Statistical analyses were conducted using Chi-square and Fisher–Freeman–Halton tests, with thematic evaluation of citation errors. Among the four LLMs evaluated, ChatGPT-4o demonstrated the highest accuracy (84%) and rationale validity (72%), followed by Gemini, Copilot, and Claude. While all models achieved perfect accuracy on lower-order (“remembering”) questions, performance declined at higher Bloom’s Taxonomy levels, though without statistically significant differences. ChatGPT-4o produced the greatest number of citations, whereas Copilot yielded the highest proportion of valid (91.8%) and relevant (48.7%) references. Despite these variations, all models showed comparable susceptibility to hallucinations. The findings indicate that current LLMs, though promising, cannot yet substitute for human reasoning or verification, underscoring the necessity of expert oversight and robust validation mechanisms to ensure reliability and mitigate misinformation in evidence-based learning environments.


Paper Information

Conference: SEACE2026
Stream: Implementation & Assessment of Innovative Technologies in Education

The full paper is not available for this title


Virtual Presentation


Comments & Feedback

Place a comment using your LinkedIn profile

Comments

Share on activity feed

Powered by WP LinkPress

Share this Research

Posted by James Alexander Gordon