BigCodeBench
收藏资源简介:
BigCodeBench是由蒙纳士大学等机构创建的一个编程任务基准数据集,包含1140个细粒度的编程任务,涉及139个库和7个领域。该数据集通过LLM和人类专家的合作构建,旨在评估大型语言模型在解决复杂和实际编程任务中的能力。数据集内容丰富,覆盖多种编程场景,如代码完成和指令到代码的转换。创建过程中,数据集通过严格的测试案例和环境设置来验证程序行为,确保数据质量。BigCodeBench主要用于评估和推动自动化软件工程领域的发展,特别是在代码生成和复杂指令理解方面。
BigCodeBench is a programming task benchmark dataset developed by Monash University and other institutions. It comprises 1140 fine-grained programming tasks covering 139 libraries and spanning 7 domains. Constructed through collaboration between large language models (LLMs) and human experts, this dataset is designed to evaluate the capabilities of LLMs in solving complex and practical programming tasks. Boasting rich content, it covers diverse programming scenarios such as code completion and instruction-to-code translation. During its development, rigorous test cases and environment configurations are employed to validate program behavior and guarantee data quality. BigCodeBench is primarily utilized to assess and promote the advancement of automated software engineering, especially in the fields of code generation and complex instruction understanding.




