mukunda1729/hallucination-risk-cases
收藏资源简介:
该数据集名为幻觉风险案例,包含20个手工标记的(提示→响应→真实情况)三元组,覆盖了大型语言模型(LLM)常见的幻觉失败模式。每个案例都根据幻觉风险(低、中、高)进行了评级,旨在帮助用户评估其幻觉检测器、评分器或判断器是否能正确区分安全响应和虚构响应。数据集包含多种类别,如事实性、虚构引用、虚构API、虚构地点、虚构事件、虚构事实、虚构作品、虚构引用、算术、统计、摘要、技术性、否定声明和未来事件等,每个类别都有相应的案例数量和测试内容。数据集采用JSON格式存储,包含ID、提示、响应、真实情况、幻觉风险等级、类别和注释等字段。
The dataset is named hallucination-risk-cases and contains 20 hand-labeled (prompt → response → ground-truth) tuples covering common LLM hallucination failure modes. Each case is rated for hallucination risk (low, medium, high) so you can evaluate whether your detector / scorer / judge correctly distinguishes the safe responses from the fabricated ones. The dataset includes various categories such as factual, fabricated-citation, fabricated-api, fabricated-place, fabricated-event, fabricated-fact, fabricated-work, fabricated-quote, arithmetic, statistical, summary, technical, negative-claim, and future-event, each with corresponding case counts and test purposes. The data is stored in JSON format, including fields such as ID, prompt, response, ground_truth, hallucination_risk, category, and notes.



