XhotpotQA-GLM52-Judge-V1
收藏资源简介:
XHotpotQA GLM-5.2翻译判断数据集V1是一个用于跨语言多跳问答翻译质量审计的标注数据集。该数据集包含2,760个评分单元,覆盖23种语言(阿拉伯语、孟加拉语、德语、希腊语、西班牙语、波斯语、法语、印地语、印尼语、意大利语、日语、韩语、荷兰语、波兰语、葡萄牙语、俄语、瑞典语、斯瓦希里语、泰语、土耳其语、乌尔都语、越南语、中文普通话)。每个单元由GLM-5.2大语言模型作为裁判,对从XHotpotQA数据集中提取的英文源文本(段落、问题或简短答案)的翻译进行0-100分的评分。数据集采用语言平衡的确定性抽样策略,每语言包含120个单元(80个段落、20个问题、20个答案)。评分基于忠实度(60%)、术语/实体/数字(15%)、流畅度(20%)和风格/语域(5%)的加权准则,并设有关键错误下限。该审计快照为V1版本,与独立的V2版本为无配对比较设计。公开架构包含judge_record_id、instance_id、target_language、unit、score等21个字段,但不包含源文本或候选译文内容,仅提供SHA-256哈希用于验证连接。数据集的加权平均得分为92.097,各单元和语言层级的详细结果在数据集中以表格形式提供。该数据集适用于研究LLM作为翻译裁判的可靠性、跨语言翻译质量评估以及多跳QA系统的翻译效果分析。
XHotpotQA GLM-5.2 Translation Judgment Dataset V1 is an annotated dataset for auditing translation quality in cross-lingual multi-hop QA. It contains 2,760 scoring units covering 23 languages (Arabic, Bengali, German, Greek, Spanish, Persian, French, Hindi, Indonesian, Italian, Japanese, Korean, Dutch, Polish, Portuguese, Russian, Swedish, Swahili, Thai, Turkish, Urdu, Vietnamese, Mandarin Chinese). Each unit is scored by the GLM-5.2 large language model on a 0-100 scale for translations of English source texts (paragraphs, questions, or short answers) extracted from the XHotpotQA dataset. The dataset uses a language-balanced deterministic sampling strategy, with 120 units per language (80 paragraphs, 20 questions, 20 answers). Scoring is based on weighted criteria: fidelity (60%), terminology/entities/numbers (15%), fluency (20%), and style/register (5%), with a critical error floor. This snapshot is V1, designed for unpaired comparison with the independent V2 version. The public schema includes 21 fields such as judge_record_id, instance_id, target_language, unit, score, etc., but does not include source text or candidate translations, only SHA-256 hashes for verification. The weighted average score is 92.097, with detailed results at unit and language levels provided in tables. The dataset is suitable for studying the reliability of LLMs as translation judges, cross-lingual translation quality assessment, and translation effect analysis of multi-hop QA systems.
XHotpotQA GLM-5.2 翻译质量审计数据集(V1)
数据集概述
这是 XHotpotQA 资源分析中使用的 V1 翻译质量审计数据集,包含 2,760 条由 LLM(请求模型别名为 glm-5.2)生成的审计标注,对翻译质量进行 0-100 分评分。数据集不包含 QA 系统预测结果或人工黄金标签,而是聚焦于翻译质量的独立审计。公共数据已去除隐藏推理过程和原始文本,通过稳定的 XHotpotQA 标识符及 UTF-8 SHA-256 哈希实现可复现连接。
数据集规模与结构
- 审计单元总数: 2,760 条,全部成功评分(100%)
- 目标语言: 23 种,采用均衡抽样
- 单元类型分布(每种语言):
- 段落: 80 条(总计 1,840)
- 问题: 20 条(总计 460)
- 答案: 20 条(总计 460)
- 加权平均分: 92.097(0-100 评分尺度)
- 抽样种子: 20260810
- 数据来源: 训练集衍生 1,866 条,验证集衍生 894 条
评分结果
按单元类型统计
| 单元 | 数量 | 均值 | 中位数 | 标准差 | <60 | <80 | ≥90 |
|---|---|---|---|---|---|---|---|
| 段落 | 1,840 | 90.913 | 94 | 9.47 | 24 | 182 | 1,358 |
| 问题 | 460 | 93.459 | 96 | 10.71 | 12 | 35 | 387 |
| 答案 | 460 | 95.474 | 100 | 14.47 | 21 | 28 | 416 |
各语言总体均分(节选)
- 最高: 葡萄牙语 95.87、西班牙语 95.39、意大利语 95.31
- 最低: 土耳其语 89.18、斯瓦希里语 89.38、瑞典语 89.90
数据字段说明
主要字段包括:
judge_record_id: 版本与单元的确定性哈希instance_id,source_id,source_split: 稳定数据集标识符target_language,unit,paragraph_id: 审计层级与单元标识score: 0-100 整数评分score_origin: 评分来源解析judge_explanation: 可视短解释(若有)source_text_sha256,candidate_text_sha256: 精确连接验证哈希requested_judge_model: 请求的模型别名(glm-5.2)resolved_judge_revision: 为空(提供方解析的模型修订版未记录)judge_prompt_version,judge_prompt_sha256: 评分合约标识符
注意: 数据集中不包含 source_text、candidate_text、隐藏推理内容、端点 URL、请求错误、API 密钥或原始日志。
模型溯源
| 设置 | 值 |
|---|---|
| 请求模型字符串 | glm-5.2 |
| 提供方解析修订版 | 未记录 |
| 温度 | 0.0 |
| 种子 | 20260810 |
| 最大输出 token | 4,000 |
| 并发工作数 | 1 |
| 最终完成数 | 2,760 / 2,760 |
身份状态标记为 requested_alias_only_provider_revision_not_recorded,表明模型别名解析的检查点无法验证。
评分来源分布
| 来源 | 行数 |
|---|---|
显式 SCORE: 行 |
2,746 |
| 推理回退中的显式评分 | 4 |
| 最后整数推理回退 | 10 |
方法论与抽样设计
- 抽样为确定性且覆盖 23 种目标语言
- 段落源文本从固定的英文 HotpotQA 上下文中恢复;问题和答案来源来自原始英文字段
- 与固定版本的 XHotpotQA V1.1 数据集通过
instance_id,unit,paragraph_id连接,并使用 SHA-256 哈希验证字符串 - V1 与 V2 审计独立抽取、非配对,均值差异不能解释为实例级或因果性的 V2 改进估计
局限性
- 单一 LLM 评判不能替代双语人工标注或标注者间一致性验证
- 提供方解析的模型版本及服务器配置不可用
- 温度为零和种子设置不能保证第三方端点行为的确定性
- 评分解析器历史上允许整数回退,
score_origin字段使这些情况透明可见 - 评分标准对不同语言和实体类型可能应用转写预期不均
- 样本均衡设计,但并非所有翻译的全量普查
许可证与引用
- 许可证: CC BY-SA 4.0(衍生自 HotpotQA/XHotpotQA 内容)
- 配套构建软件: MIT 许可证
- 论文: arXiv:2608.27481,DOI: 10.48550/arXiv.2608.27481
- 固定版本: 公共注释负载冻结于 Hub 修订版
ba891ae62ed989606c9fc2fd5f08f9e88ef37547 - 相关资源:
- 代码库: https://github.com/Iman998/XhotpotQA
- 数据集合集: https://huggingface.co/collections/Iman998/xhotpotqa-cross-lingual-multi-hop-qa-6a888df6aee4a4f5612c3a1a
- 源数据集: https://huggingface.co/datasets/Iman998/XhotpotQA




