| 1 | Title of the Article | Evaluating the Efficiency of Large Language Models in Detecting and Mitigating Hallucinations Using a Curated Domain-Oriented Dataset and Post-Hoc Retrieval |
| 2 | Author's name | Chandra Sekhar Sanaboina: Assistant Professor, Department of Computer Science and Engineering, UCEK, JNTUK, Kakinada, Andhra Pradesh, India |
| 3 | Author's name | Tammisetti Surekha |
| 4 | Subject | Computer Science |
| 5 | Keyword(s) | Large Language Models, Hallucination Detection, Claim Verification, Evidence Retrieval, Hallucination Mitigation, FActScore, Semantic re?ranking, Factuality Assessment |
| 6 | Abstract | Hallucination in large language models (LLMs) refers to the generation of plausible and coherent information that is incorrect, unsupported, or inconsistent with the available evidence. Detecting such hallucinations and verifying the information in generated responses are important for improving the reliability of LLM based systems. This study evaluates the ability of LLMs to detect hallucinated responses and investigates whether external evidence can improve the verification process. A curated dataset of 700 question and answer pairs covering 20 academic disciplines was used, including both factual and hallucinated responses. Four models, GPT 3.5 Turbo, GPT 4, GPT 4 Turbo, and Gemma 7B, were first evaluated using their internal knowledge. For responses that required further checking, the responses were divided into individual claims, and relevant evidence was retrieved from Wikipedia using BM25 and semantic re ranking. Each claim was then verified against the retrieved evidence, and FActScore was used to support the final classification as factual or hallucinated. With no evidence the models scored accuracies between 0.653 and 0.943. After adding evidence from Wikipedia accuracy rose from 0.88 to 0.986. Fleiss’ Kappa also increased from 0.544 to 0.829, indicating greater agreement among the models after external evidence was introduced. These results show that relevant external evidence can improve the reliability and consistency of LLM based hallucination detection. The proposed approach uses retrieval to verify claims and reclassify responses when the initial judgment needs evidence. Retrieval does not directly correct the content. This checking step can therefore help mitigating hallucination by spotting information that is not supported. |
| 7 | Publisher | Innovative Research Publication |
| 8 | Journal Name; vol., no. | International Journal of Innovative Research in Computer Science & Technology (IJIRCST); Volume-14 Issue-5 |
| 9 | Publication Date | Sep-Oct 2026 |
| 10 | Type | Peer-reviewed Article |
| 11 | Format | |
| 12 | Uniform Resource Identifier | https://ijircst.org/view_abstract.php?title=Evaluating-the-Efficiency-of-Large-Language-Models-in-Detecting-and-Mitigating-Hallucinations-Using-a-Curated-Domain-Oriented-Dataset-and-Post-Hoc-Retrieval&year=2026&vol=14&primary=QVJULTE0Nzg= |
| 13 | Digital Object Identifier(DOI) | 10.55524/ijircst.2026.14.5.3 https://doi.org/10.55524/ijircst.2026.14.5.3 |
| 14 | Language | English |
| 15 | Page No | 19-33 |