遇见数据集

nygdon/vimqa-generated-answers-pass1

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是Vi-MQA数据集(来自VMLU Benchmark Suite)的Pass 1生成答案和评估结果,包含总计4,762条记录。数据集主要用于越南语的多选问答(MQA)任务,涉及三个语言模型(Gemma 4 31B IT、Llama 4 Scout、Qwen3 32B)的输出比较。评估结果分为三个部分:agree.jsonl(1,152个样本,所有三个模型给出完全相同的答案,置信度高)、majority.jsonl(381个样本,三个模型中有两个一致,可靠性良好)和conflict.jsonl(46个样本,无共识,需要进一步的Pass 2和手动审查)。数据集还包括原始模型输出文件和评估脚本,用于分析模型在VMLU基准测试中的表现和一致性。

This repo contains the Pass 1 outputs and evaluation results for the Vi-MQA Dataset from the VMLU Benchmark Suite with a total of 4,762 records. It includes formatted outputs from three models (Gemma 4 31B IT, Llama 4 Scout, Qwen3 32B) and evaluation results categorized into agree.jsonl (1,152 samples where all three models give the exact same answer), majority.jsonl (381 samples where two out of three models agree), and conflict.jsonl (46 samples with no consensus requiring Pass 2 and manual review). The dataset is designed for Vietnamese multiple-choice question answering (MQA) tasks within the VMLU benchmark framework.

提供机构:
nygdon
二维码
社区交流群
二维码
科研交流群
商业服务