遇见数据集

ATH-MaaS/Marco-Bench-MIF

收藏
Hugging Face2025-08-14 更新2026-07-22 收录
官方服务:

资源简介:

--- license: apache-2.0 --- # Marco-Bench-MIF: A Benchmark for Multilingual Instruction-Following Evaluation [![License: Apache 2.0](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://www.apache.org/licenses/LICENSE-2.0) [![ACL 2025](https://img.shields.io/badge/ACL-2025-blue)](https://aclanthology.org/2025.acl-long.1172/) [![arXiv](https://img.shields.io/badge/arXiv-2507.11882-b31b1b.svg)](https://arxiv.org/abs/2507.11882) ## Introduction Marco-Bench-MIF is the first deeply localized multilingual benchmark designed to evaluate instruction-following capabilities across 30 languages. Unlike existing benchmarks that rely primarily on machine translation, Marco-Bench-MIF implements fine-grained cultural adaptations to provide more accurate assessment. Our research demonstrates that machine-translated data underestimates model performance by 7-22% in multilingual environments. ## Key Features - **Extensive Language Coverage**: 30 languages spanning 6 major language families, including high-resource (English, Chinese, German) and low-resource languages (Yoruba, Nepali) - **Deep Cultural Localization**: Three-step process of lexical replacement, theme transformation, and pragmatic reconstruction to ensure cultural and linguistic appropriateness - **Diverse Constraint Types**: 541 instruction-response pairs covering single/multiple constraints, expressive/content constraints, and various instruction types - **Comparative Dataset**: Machine-translated and culturally-localized versions available for specific languages (Arabic, Chinese, Spanish, etc.) to enable comparative research ## Dataset Access The dataset will be available through our GitHub repository and Hugging Face: ```bash git clone https://github.com/AIDC-AI/Marco-Bench-MIF.git ``` ## Key Findings Our benchmark evaluated 20+ LLM models and revealed: 1. Model scale strongly correlates with performance, with 70B+ models outperforming 8B models by 45-60% 2. A 25-35% performance gap exists between high-resource languages (German, Chinese) and low-resource languages (Yoruba, Nepali) 3. Significant differences between localized and machine-translated evaluations, especially for complex instructions ## Contact For questions or suggestions, please submit a GitHub issue or contact us: - Email: lyuchenyang.lcy@alibaba-inc.com - Project homepage: https://github.com/AIDC-AI/Marco-Bench-MIF ## License This dataset is licensed under the [Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0). ## Acknowledgments Special thanks to all annotators and translators who participated in dataset construction and validation. This project is supported by Alibaba International Digital Commerce Group.

--- 许可证:Apache 2.0 --- # Marco-Bench-MIF:多语言指令遵循评估基准 [![许可证:Apache 2.0](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://www.apache.org/licenses/LICENSE-2.0) [![ACL 2025](https://img.shields.io/badge/ACL-2025-blue)](https://aclanthology.org/2025.acl-long.1172/) [![arXiv](https://img.shields.io/badge/arXiv-2507.11882-b31b1b.svg)](https://arxiv.org/abs/2507.11882) ## 引言 Marco-Bench-MIF是首个经过深度本土化处理的多语言基准测试集,用于评估跨30种语言的指令遵循能力。与当前主要依赖机器翻译的现有基准不同,Marco-Bench-MIF采用细粒度的文化适配方案,以实现更精准的性能评估。本研究表明,在多语言场景中,机器翻译数据会低估模型7%-22%的实际性能。 ## 核心特性 - **广泛的语言覆盖范围**:涵盖6大语系的30种语言,包含高资源语言(英语、中文、德语)与低资源语言(约鲁巴语、尼泊尔语) - **深度文化本土化**:通过词汇替换、主题重构、语用重建三步流程,确保内容符合文化与语言规范 - **多样化约束类型**:包含541组指令-响应对,覆盖单约束/多约束、表达型/内容型约束以及多种指令类型 - **可对比数据集**:针对阿拉伯语、中文、西班牙语等特定语言,提供机器翻译版与文化适配版数据,支持对比研究 ## 数据集获取 本数据集将通过GitHub仓库与Hugging Face平台发布: bash git clone https://github.com/AIDC-AI/Marco-Bench-MIF.git ## 核心发现 本基准对20余款大语言模型(Large Language Model,LLM)进行了评估,结果显示: 1. 模型规模与性能呈强正相关,70B参数及以上的模型较8B参数模型性能高出45%-60% 2. 高资源语言(德语、中文)与低资源语言(约鲁巴语、尼泊尔语)之间存在25%-35%的性能差距 3. 文化适配版评估与机器翻译版评估之间存在显著差异,在复杂指令场景中尤为明显 ## 联系方式 如有疑问或建议,请提交GitHub Issue或联系我们: - 邮箱:lyuchenyang.lcy@alibaba-inc.com - 项目主页:https://github.com/AIDC-AI/Marco-Bench-MIF ## 许可证 本数据集采用[Apache 2.0许可证](https://www.apache.org/licenses/LICENSE-2.0)进行授权。 ## 致谢 特别感谢所有参与数据集构建与验证的标注人员与翻译人员。本项目得到阿里巴巴国际数字商业集团的支持。

提供机构:
ATH-MaaS
二维码
社区交流群
二维码
科研交流群
商业服务