遇见数据集

SUS LLM Hallucination Dataset: A curated benchmark for evaluating hallucinations in large language models in a regulated public health system

收藏
Zenodo2026-03-21 更新2026-05-26 收录
官方服务:

资源简介:

This dataset provides a curated benchmark for evaluating hallucinations in large language models (LLMs) in a highly regulated public health system. The dataset contains 90 structured questions covering key domains of Brazil’s Unified Health System (SUS), including coverage, care pathways, eligibility criteria, and patient rights. Each item is annotated with a taxonomy of hallucination types (T1–T6), risk levels, and associated normative sources, enabling systematic evaluation of factuality, compliance, and safety in LLM-generated responses. The dataset is designed for use in:- Benchmarking LLM performance in regulated environments- Evaluating retrieval-augmented generation (RAG) systems- Research in AI safety, governance, and hallucination mitigation- Public sector AI evaluation and policy-oriented research This resource emphasizes high-risk, regulation-sensitive scenarios, where hallucinations may have real-world consequences. All data are provided in structured formats with accompanying documentation to support reproducibility and reuse.

提供机构:
Zenodo
创建时间:
2026-03-21
二维码
社区交流群
二维码
科研交流群
商业服务