证据数据集(Evidence Dataset)
收藏资源简介:
该数据集是为了研究语言模型在面对不同类型和可靠性的证据时的置信度和响应变化。数据集包含从SciQ、TriviaQA和GSM8K数据集中生成的各种证据类型,例如与问题一致的黄金证据、包含错误信息的冲突证据、不完整证据、矛盾证据、不相关证据和偶然证据。数据集还考虑了证据的可靠性因素,如来源的可信度、详细程度、时效性和实验性。该数据集旨在帮助理解为什么LLMs偏离贝叶斯认识论,并评估LLMs在处理不同类型和强度证据时的表现。
This dataset is designed to study the confidence and response variations of language models when exposed to evidence of different types and reliability. It includes various types of evidence generated from the SciQ, TriviaQA, and GSM8K datasets, such as golden evidence consistent with the question, conflicting evidence containing misinformation, incomplete evidence, contradictory evidence, irrelevant evidence, and coincidental evidence. The dataset also accounts for reliability factors of evidence, including source credibility, level of detail, timeliness, and experimental nature. This dataset aims to help understand why LLMs deviate from Bayesian epistemology, and evaluate the performance of LLMs when processing evidence of different types and strengths.
数据集概述
基本信息
- 数据集名称:From Evidence to Belief: A Bayesian Epistemology Approach to Language Models
- 会议信息:NAACL 2025, main
相关说明
- 该数据集与贝叶斯认识论方法在语言模型中的应用相关。




