mHumanEval
收藏资源简介:
mHumanEval是由乔治梅森大学开发的用于评估大型语言模型代码生成能力的多语言基准数据集。该数据集包含33,456个编程提示,涵盖204种自然语言和25种编程语言,旨在解决现有基准在任务多样性、测试覆盖率和语言范围方面的局限性。数据集通过机器翻译和专家人工翻译相结合的方式创建,确保了翻译质量。mHumanEval的应用领域主要集中在评估和提升多语言环境下的代码生成模型性能,特别是在低资源语言环境中的表现。
mHumanEval is a multilingual benchmark dataset developed by George Mason University for evaluating the code generation capabilities of large language models. This dataset includes 33,456 programming prompts, covering 204 natural languages and 25 programming languages, and aims to address the limitations of existing benchmarks in terms of task diversity, test coverage, and language scope. The dataset is constructed through a combination of machine translation and expert manual translation to ensure translation quality. The main application areas of mHumanEval focus on evaluating and enhancing the performance of code generation models in multilingual environments, especially their performance in low-resource language settings.

- 1mHumanEval -- A Multilingual Benchmark to Evaluate Large Language Models for Code Generation乔治梅森大学 · 2024年



