Global PIQA
收藏资源简介:
Global PIQA是一个包含100多种语言的参与式常识推理基准数据集,由来自65个国家的335名研究人员手工构建。该数据集涵盖116种语言变体,包含五大洲、14个语系和23种文字系统。在非平行分割版本中,超过50%的示例涉及当地食物、习俗、传统或其他文化特定元素。每个示例包含一个提示和两个候选解决方案,一个正确一个错误,需要物理常识推理能力进行判断。
Global PIQA is a participatory commonsense reasoning benchmark dataset covering over 100 languages, manually constructed by 335 researchers from 65 countries. It encompasses 116 language variants, spanning five continents, 14 language families, and 23 writing systems. In its non-parallel split version, over 50% of the examples involve local food, customs, traditions, or other culture-specific elements. Each example contains a prompt and two candidate solutions—one correct and one incorrect, which require physical commonsense reasoning ability for judgment.
Global PIQA 数据集概述
数据集基本信息
- 数据集名称:Global PIQA v0.1
- 构建方式:由来自65个国家的335名研究人员手工构建的参与式常识推理基准
- 语言覆盖:涵盖100多种语言,包含116种语言变体,覆盖五大洲、14个语系和23种文字系统
核心特征
- 数据格式:每个示例包含一个提示和两个候选解决方案(一个正确、一个错误)
- 推理类型:需要物理常识推理,包括物体物理属性、功能、物理和时间关系及日常活动知识
- 文化特色:在非平行分割中,超过50%的示例涉及当地食物、习俗、传统或其他文化特定元素
数据集用途
- 主要用途:大型语言模型评估
- 附加价值:展示人类语言所嵌入的广泛文化多样性
获取方式
- 数据集地址:https://huggingface.co/datasets/mrlbenchmarks/global-piqa-nonparallel
- 论文链接:https://arxiv.org/abs/2510.24081
许可信息
- 许可证:CC BY-SA 4.0
- 使用限制:禁止用于AI系统训练或作为合成数据种子,仅限LLM评估用途
版本计划
- 未来版本:Global PIQA v1计划扩展语言覆盖范围并添加平行分割数据集
引用信息
bibtex @article{mrl-workshop-2025-global-piqa, title={Global {PIQA}: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures}, author={Tyler A. Chang et al.}, journal={Preprint}, year={2025}, url={https://arxiv.org/abs/2510.24081} }




