atreydesai/augmented-mcqa-gemini-augmented
收藏资源简介:
该数据集是一个用于评估和比较人类与AI模型(如Google Gemini 3.1 Pro预览版)在问答任务中表现的数据集。它包含2423个训练示例,每个示例具有唯一标识符(id、sample_id、question_id)、问题文本、答案、类别、选项列表,以及多组对比数据:人类从头生成的回答、模型从头生成的回答、数据增强后的人类回答、数据增强后的模型回答,以及消融实验的增强数据。这些数据还包含了随机化选项和正确答案字母标识,旨在支持对问答性能、数据增强效果和模型行为的分析研究。
This dataset is designed for evaluating and comparing the performance of humans and AI models (such as Google Gemini 3.1 Pro preview) on question-answering tasks. It contains 2423 training examples, each with unique identifiers (id, sample_id, question_id), question text, answer, category, option lists, and multiple comparative data groups: human-from-scratch responses, model-from-scratch responses, augmented human responses, augmented model responses, and ablation-augmented data. The dataset also includes randomized options and correct answer letter indicators, aiming to support analysis of QA performance, data augmentation effects, and model behavior.




