gpqa-ptpt
收藏资源简介:
GPQA-PT数据集是Graduate-Level Google-Proof Q&A Benchmark(GPQA)的葡萄牙语机器翻译版本。GPQA是一个具有挑战性的专家级问答基准数据集,包含需要深入领域知识才能回答的高难度问题。该翻译版本使用经过微调的GemmaX2-9B模型将原始英文数据集翻译为欧洲葡萄牙语,但仅翻译了问题、正确答案、错误答案和解释这几个核心字段。数据集包含四个配置:gpqa_extended、gpqa_main、gpqa_diamond和gpqa_experts,每个配置都提供训练集分割。作为AMALIA项目的一部分,该数据集被纳入AMALIA-Bench基准套件,专门用于评估大型语言模型在欧洲葡萄牙语上的能力。由于是机器翻译生成,数据集可能包含翻译错误或语言伪影。数据集适用于问答和文本生成任务,特别适合作为专家级知识基准测试,用于衡量模型在复杂领域问题上的理解和推理能力。
The GPQA-PT dataset is the European Portuguese machine translation variant of the Graduate-Level Google-Proof Q&A Benchmark (GPQA). GPQA is a challenging expert-level question answering benchmark dataset comprising high-difficulty questions that demand in-depth domain expertise to answer. This translated iteration uses the fine-tuned GemmaX2-9B model to render the original English dataset into European Portuguese, with only core fields including questions, correct answers, distractors, and explanations translated. The dataset includes four configurations: gpqa_extended, gpqa_main, gpqa_diamond, and gpqa_experts, each of which provides a training set split. As a component of the AMALIA project, this dataset is integrated into the AMALIA-Bench benchmark suite, which is dedicated to evaluating the capabilities of large language models on European Portuguese tasks. Given that it is generated via machine translation, the dataset may contain translation errors or linguistic artifacts. The dataset is applicable to question answering and text generation tasks, and is particularly suitable as an expert-level knowledge benchmark for measuring models' understanding and reasoning abilities on complex domain-specific questions.
数据集概述
- 数据集名称: GPQA-PT
- 数据集描述: 这是一个研究生级别的、谷歌难以解答的问答基准的葡萄牙语机器翻译版本,包含专家级问题。
- 语言: 葡萄牙语 (pt)
- 任务类别: 问答、文本生成
- 标签: 基准、专家级、amalia-bench
数据集配置
该数据集包含四个子集,每个子集均为训练集:
- gpqa_extended: 数据文件位于
gpqa_extended/train-* - gpqa_main: 数据文件位于
gpqa_main/train-* - gpqa_diamond: 数据文件位于
gpqa_diamond/train-* - gpqa_experts: 数据文件位于
gpqa_experts/train-*
翻译说明
- 仅有以下列被翻译:问题、正确答案、错误答案和解释。
- 翻译模型为针对欧洲葡萄牙语微调的 GemmaX2-9B。
- 警告: 该数据集由机器翻译生成,可能包含翻译错误或人工产物。
原始数据集
原始数据集为 Idavidrein/gpqa。
归属项目
该数据集是 AMALIA 项目 的一部分,并收录于 AMALIA-Bench,这是一个用于评估欧洲葡萄牙语大语言模型的综合基准套件。





