CS-Bench - 计算机科学领域大型语言模型综合基准测试数据集
收藏资源简介:
CS-Bench由北京邮电大学构建,是首个致力于评估大型语言模型(LLMs)在计算机科学领域表现的双语(中英)基准测试数据集。该数据集包含约5000个精心策划的测试样本,覆盖计算机科学的4个主要领域及26个子领域,包含多种任务形式和知识推理类型。数据集的内容涵盖了计算机科学领域的广泛主题,包括但不限于编程语言、算法、数据结构等。通过CS-Bench,研究人员对30多个主流大型语言模型进行了全面评估,揭示了模型规模与计算机科学表现之间的关系,并定量分析了现有模型的失败原因,指出了改进方向,包括知识补充和特定于计算机科学的推理能力。
CS-Bench, constructed by Beijing University of Posts and Telecommunications, is the first bilingual (Chinese-English) benchmark dataset dedicated to evaluating the performance of large language models (LLMs) in the field of computer science. The dataset comprises approximately 5,000 meticulously curated test samples, covering 4 major areas and 26 subfields of computer science, and includes a variety of task formats and knowledge reasoning types. The content of the dataset spans a wide range of topics in computer science, including but not limited to programming languages, algorithms, and data structures. Through CS-Bench, researchers have conducted a comprehensive evaluation of over 30 mainstream large language models, revealing the relationship between model scale and performance in computer science, quantitatively analyzing the failure reasons of existing models, and pointing out directions for improvement, including knowledge supplementation and computer science-specific reasoning capabilities.
CS-Bench数据集概述
数据集名称
- 名称: CS-Bench
- 全称: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery
数据集目的
- 目的: 评估大型语言模型在计算机科学领域的性能,涵盖26个子领域,包括数据结构与算法、计算机组织、计算机网络和操作系统等。
数据集构成
- 样本数量: 约5000个精心策划的测试样本
- 覆盖领域: 4个关键领域,26个子领域
- 任务类型: 知识型和推理型任务
数据集详细信息
- 语言: 中英双语
- 详细统计:
- 问题与答案长度分布: 提供英文和中文的问题与答案长度分布图
- 子领域详情: 包含26个子领域的详细分类和示例
数据集使用
- 评估模型: 已对超过30个主流大型语言模型进行评估
- 评估结果: 提供详细的模型性能评估和分析,包括知识补充和特定领域推理的改进方向
引用信息
- 引用格式: latex @article{song2024cs, title={CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery}, author={Song, Xiaoshuai and Diao, Muxi and Dong, Guanting and Wang, Zhengyang and Fu, Yujia and Qiao, Runqi and Wang, Zhexu and Fu, Dayuan and Wu, Huangxuan and Liang, Bin and others}, journal={arXiv preprint arXiv:2406.08587}, year={2024} }
数据集链接
- Huggingface链接: CS-Bench on Huggingface




