MAQA* and AmbigQA*
收藏资源简介:
MAQA*和AmbigQA*是首批配备有来自事实共现估计的真实答案分布的模糊问答(QA)数据集。这些数据集首次允许在现实世界的模糊条件下对不确定性估计器进行原理性评估。数据集的内容包括显式地面真实答案分布p*,这些分布是从事实共现统计中估计的。数据集创建过程涉及到收集模糊的问答对,并估计每个问题的真实答案分布。这些数据集旨在解决当前LLMs不确定性量化方法在实际应用中的不足,特别是在处理具有非平凡随机性的问题时。
MAQA* and AmbigQA* are the first ambiguous question answering (QA) datasets equipped with ground-truth answer distributions derived from factual co-occurrence estimation. These datasets enable the first principled evaluation of uncertainty estimators under real-world ambiguous conditions. The datasets contain explicit ground-truth answer distributions p*, which are estimated from factual co-occurrence statistics. The dataset creation process involves collecting ambiguous QA pairs and estimating the ground-truth answer distribution for each question. These datasets are designed to address the limitations of current uncertainty quantification methods for large language models (LLMs) in practical applications, particularly when handling questions with nontrivial randomness.

- 1The Illusion of Certainty: Uncertainty quantification for LLMs fails under ambiguity慕尼黑工业大学计算、信息和科技学院 & 慕尼黑数据科学研究所 · 2025年



