LoCoBench
收藏资源简介:
LoCoBench 是一个为评估长上下文大型语言模型在复杂软件工程中的应用而设计的综合基准。该数据集由 Salesforce AI Research 创建,包含 8000 个评估场景,覆盖 10 种编程语言和 36 个领域类别。数据集的上下文长度从 10K 到 1M Tokens 不等,能够精确评估长上下文性能的退化。LoCoBench 引入了 8 个任务类别,包括架构理解、跨文件重构、多会话开发、错误调查、功能实现、代码理解、集成测试和安全分析,旨在解决复杂软件工程中的长上下文能力评估问题。数据集通过一个五阶段的流程创建,包括项目规范生成、代码库生成、评估场景创建、验证和质量管理以及 LLM 评估和评分。LoCoBench 提供了一个全面的评估框架,包含 17 个指标,涵盖软件工程卓越、功能正确性、代码质量评估和长上下文利用等方面。
LoCoBench is a comprehensive benchmark designed to evaluate the applications of long-context large language models in complex software engineering. Developed by Salesforce AI Research, this dataset comprises 8,000 evaluation scenarios spanning 10 programming languages and 36 domain categories. The context lengths of the dataset range from 10K to 1M Tokens, enabling precise evaluation of performance degradation in long-context scenarios. LoCoBench introduces 8 task categories, including architecture understanding, cross-file refactoring, multi-session development, bug investigation, function implementation, code comprehension, integration testing, and security analysis, aiming to address the challenge of evaluating long-context capabilities in complex software engineering. The dataset is created through a five-stage workflow, which includes project specification generation, codebase generation, evaluation scenario creation, validation and quality management, as well as LLM evaluation and scoring. LoCoBench provides a comprehensive evaluation framework with 17 metrics covering software engineering excellence, functional correctness, code quality assessment, and long-context utilization, among other aspects.

- 1通过Salesforce AI Research · 2025年



