遇见数据集

Benchmark Dataset for Evaluating Hallucination Handling and Factual Robustness in Large Language Models

收藏
Zenodo2026-06-12 更新2026-06-18 收录
官方服务:

资源简介:

This repository contains the evaluation materials used to assess the factual robustness and hallucination-handling capabilities of Large Language Models (LLMs). The benchmark includes a collection of manually designed test cases covering different challenging situations, such as: Common factual questions. Less frequent historical facts. Conceptual error correction. Detection of false premises. Recognition of nonexistent entities. Identification of fictional events and characters. Historical misconceptions and factual inconsistencies. For each test case, the repository provides: The reference answer (ground truth). Model-generated responses from multiple executions. Semantic similarity scores computed using SBERT. Quality assessment scores computed using BLEURT. Response latency measurements. Aggregated evaluation statistics. The dataset was created to support reproducible research on factual reliability, hallucination detection, and educational applications of generative AI systems. Researchers are encouraged to reuse, extend, and compare their models against this benchmark under the terms of the accompanying license.

提供机构:
Zenodo
创建时间:
2026-06-12
二维码
社区交流群
二维码
科研交流群
商业服务