Sandhya1912/AA-Omniscience-Public
收藏资源简介:
AA-Omniscience-Public 是一个基准数据集,包含600个跨多个领域的问题,用于测试大型语言模型的知识可靠性和幻觉倾向。该数据集旨在衡量模型在跨领域准确回忆事实信息的能力,以及在知识不足时正确弃权的能力。完整数据集包含6000个问题,覆盖经济重要领域,而公共版本是完整数据集的10%代表性子集,以确保评估完整性。问题通过自动化问题生成代理创建,基于权威来源,并经过相似性、难度和模糊性过滤,具有难度高、答案明确、不依赖特定来源和答案精确的特点。数据集用于评估模型的知识准确性和幻觉率,但公共版本在领域或主题级别的结果可能不可靠,因为样本量较小。
AA-Omniscience-Public contains 600 questions across a wide range of domains used to test a model’s knowledge and hallucination tendencies. It is a benchmark dataset designed to measure a model’s ability to both recall factual information accurately across domains, and correctly abstain when its knowledge is insufficient. The full dataset comprises 6,000 total questions, split across economically significant domains, and the public version is a 10% subset of the full question set, sampled to be representative. Questions are created using a question generation agent, derived from authoritative sources and filtered based on similarity, difficulty, and ambiguity, making them difficult, unambiguous, not reliant on specific sources, and precise. The dataset is used to evaluate model knowledge accuracy and hallucination rates, though the public set may not be reliable at a domain or topic level due to small size.




