CS-Bench
收藏资源简介:
CS-Bench是首个专注于评估大型语言模型在计算机科学领域表现的双语(中英文)基准。该数据集包含约5000个精心筛选的测试样本,覆盖计算机科学的26个子领域,涉及多种任务形式和知识与推理的划分。CS-Bench不仅评估模型对计算机科学知识的掌握,还评估其应用这些知识进行推理的能力。此外,支持中英文双语评估,使得能够跨语言环境评价模型的性能。该数据集旨在解决当前大型语言模型在计算机科学领域评估不足的问题,推动模型在教育、工业和科学等领域的应用。
CS-Bench is the first bilingual (Chinese and English) benchmark specifically focused on evaluating the performance of large language models (LLMs) in the field of computer science. This dataset contains approximately 5,000 carefully curated test samples, covering 26 subfields of computer science, and involves a variety of task formats as well as the division of evaluation into knowledge mastery and reasoning ability. CS-Bench not only evaluates a model’s mastery of computer science knowledge, but also its ability to apply such knowledge for reasoning. Additionally, it supports bilingual evaluation in both Chinese and English, enabling cross-lingual assessment of model performance. This dataset aims to address the current insufficient evaluation of large language models in the computer science domain, and promote the application of these models in fields such as education, industry, and scientific research.




