CodeAssistBench (CAB)
收藏资源简介:
CodeAssistBench (CAB) 是一个基于真实世界开发者场景的多轮编程辅助评估框架。该数据集由 3,286 个真实编程问题组成,涵盖了 231 个代码仓库,跨越了七种编程语言和多样化的问题领域。数据集的创建过程完全自动化,无需手动筛选,确保了数据集的多样性和高质量。CAB 的设计旨在解决现有编程问答基准的局限性,如单轮交互、需要大量手动筛选和无法代表真实项目环境等问题。该数据集用于评估大型语言模型在复杂、项目特定环境下的编程辅助能力。
CodeAssistBench (CAB) is a multi-turn programming assistance evaluation framework based on real-world developer scenarios. This dataset consists of 3,286 real programming problems, covering 231 code repositories, spanning seven programming languages and diverse problem domains. The dataset was created fully automatically without manual filtering, ensuring its diversity and high quality. CAB is designed to address the limitations of existing programming question-and-answer benchmarks, such as single-turn interaction, the need for extensive manual filtering, and failure to represent real-world project environments. This dataset is used to evaluate the programming assistance capabilities of large language models in complex, project-specific environments.




