International Journal of Innovative Research in Computer Science and Technology
Year: 2026, Volume: 14, Issue: 5
First page : ( 19) Last page : ( 33)
Online ISSN : 2347-5552
DOI: 10.55524/ijircst.2026.14.5.3 |
DOI URL: https://doi.org/10.55524/ijircst.2026.14.5.3
This is an Open Access article distributed under the terms of the Creative Commons Attribution License (CC BY 4.0)http://creativecommons.org/licenses/by/4.0
Article Tools: Print the Abstract | Indexing metadata | How to cite item | Email this article | Post a Comment
Chandra Sekhar Sanaboina , Tammisetti Surekha
Hallucination in large language models (LLMs) refers to the generation of plausible and coherent information that is incorrect, unsupported, or inconsistent with the available evidence. Detecting such hallucinations and verifying the information in generated responses are important for improving the reliability of LLM based systems. This study evaluates the ability of LLMs to detect hallucinated responses and investigates whether external evidence can improve the verification process. A curated dataset of 700 question and answer pairs covering 20 academic disciplines was used, including both factual and hallucinated responses. Four models, GPT 3.5 Turbo, GPT 4, GPT 4 Turbo, and Gemma 7B, were first evaluated using their internal knowledge. For responses that required further checking, the responses were divided into individual claims, and relevant evidence was retrieved from Wikipedia using BM25 and semantic re ranking. Each claim was then verified against the retrieved evidence, and FActScore was used to support the final classification as factual or hallucinated. With no evidence the models scored accuracies between 0.653 and 0.943. After adding evidence from Wikipedia accuracy rose from 0.88 to 0.986. Fleiss’ Kappa also increased from 0.544 to 0.829, indicating greater agreement among the models after external evidence was introduced. These results show that relevant external evidence can improve the reliability and consistency of LLM based hallucination detection. The proposed approach uses retrieval to verify claims and reclassify responses when the initial judgment needs evidence. Retrieval does not directly correct the content. This checking step can therefore help mitigating hallucination by spotting information that is not supported.
Assistant Professor, Department of Computer Science and Engineering, UCEK, JNTUK, Kakinada, Andhra Pradesh, India
No. of Downloads: 2 | No. of Views: 20
Arijit Das.
Sep-Oct 2026 - Vol 14, Issue 5
Madhav A. Kankhar, C Namrata Mahender.
July 2026 - Vol 14, Issue 4
Mohammad Shafeeq, Monika Tripathi.
July 2026 - Vol 14, Issue 4
