遇见数据集

A Tale of Two Hallucinations: Empirical Evidence for Distinct Failure Modes in Large Language Models Dataset

收藏
Zenodo2026-05-20 更新2026-05-26 收录
官方服务:

资源简介:

This dataset is a comprehensive test suite designed to empirically validate and evaluate two distinct types of AI hallucinations: Type 1 (Uncertainty-driven) and Type 2 (Confidence-driven). The data was collected from controlled experiments involving six model families (including Perplexity, Claude Sonnet 4, ChatGPT 5, Deepseek, ChatGPT 4o, and Gemini 2.5 Pro) and a total of 120 test cases.

提供机构:
Zenodo
创建时间:
2025-09-20
二维码
社区交流群
二维码
科研交流群
商业服务