Benchmark Dataset for Evaluating Hallucination Handling and Factual Robustness in Large Language Models
收藏资源简介:
This repository contains the evaluation materials used to assess the factual robustness and hallucination-handling capabilities of Large Language Models (LLMs). The benchmark includes a collection of manually designed test cases covering different challenging situations, such as: Common factual questions. Less frequent historical facts. Conceptual error correction. Detection of false premises. Recognition of nonexistent entities. Identification of fictional events and characters. Historical misconceptions and factual inconsistencies. For each test case, the repository provides: The reference answer (ground truth). Model-generated responses from multiple executions. Semantic similarity scores computed using SBERT. Quality assessment scores computed using BLEURT. Response latency measurements. Aggregated evaluation statistics. The dataset was created to support reproducible research on factual reliability, hallucination detection, and educational applications of generative AI systems. Researchers are encouraged to reuse, extend, and compare their models against this benchmark under the terms of the accompanying license.



