CSR-Bench
收藏资源简介:
CSR-Bench是一个针对计算机科学研究项目的代码仓库部署任务的评估基准。该数据集由来自GitHub的100个高质量代码仓库组成,这些仓库经过精心挑选,涵盖了自然语言处理、计算机视觉、大型语言模型、机器学习等多个领域。数据集中的README文件和目录结构为评估大型语言模型在代码部署任务中的性能提供了关键信息。CSR-Bench旨在评估LLM在理解指令手册、生成可执行命令以及解决部署过程中的错误等方面的能力。
CSR-Bench is an evaluation benchmark for code repository deployment tasks in computer science research projects. This dataset comprises 100 high-quality code repositories sourced from GitHub, which were carefully curated and cover multiple domains including natural language processing (NLP), computer vision (CV), large language models (LLMs), machine learning, and other related fields. The README files and directory structures within the dataset provide critical information for evaluating the performance of large language models in code deployment tasks. CSR-Bench aims to assess the capabilities of LLMs in understanding instruction manuals, generating executable commands, and troubleshooting errors encountered during the deployment process.




