ADVQA
收藏资源简介:
ADVQA是一个由马里兰大学创建的对抗性问答数据集,旨在通过高智能含量的样本挑战模型而非人类。该数据集包含9347条高质量、现实且具有挑战性的问题,这些问题通过人机交互过程生成,确保了问题的对抗性和区分度。数据集的创建过程涉及专家编写问题,并通过人机竞赛收集答案和模型预测,使用ADVSCORE评估问题质量。ADVQA的应用领域主要在于揭示语言模型的弱点,推动模型向人类智能水平靠拢,特别适用于评估和提升模型在复杂问题上的表现。
ADVQA is an adversarial question answering dataset developed by the University of Maryland, which is designed to challenge AI models rather than humans through samples with high intellectual complexity. It consists of 9,347 high-quality, realistic and challenging questions generated via human-machine interaction processes, ensuring the adversarial nature and discriminative ability of these questions. The dataset creation workflow involves experts drafting questions, followed by collecting answers and model predictions through human-machine competitions, and using ADVSCORE to evaluate question quality. ADVQA is primarily applied to uncover the inherent weaknesses of language models, advance models toward human-level intelligence, and is particularly well-suited for evaluating and improving model performance on complex questions.




