SecBench
收藏资源简介:
SecBench是一个全面的多维基准数据集,旨在评估大型语言模型(LLMs)在网络安全领域的能力。该数据集包括多种格式的问题(如多项选择题和简答题),涵盖不同的能力层次(知识保持和逻辑推理),并且支持多种语言(中文和英文)。数据集通过从公开资源收集高质量数据和组织网络安全问题设计竞赛构建而成,包含44,823个多项选择题和3,087个简答题。此外,数据集还使用了强大的LLMs进行数据标注和构建评分代理,以自动评估简答题。
SecBench is a comprehensive multi-dimensional benchmark dataset designed to evaluate the capabilities of Large Language Models (LLMs) in the cybersecurity domain. This dataset includes questions in various formats such as multiple-choice questions and short-answer questions, covers different capability levels including knowledge retention and logical reasoning, and supports multiple languages (Chinese and English). The dataset is constructed by collecting high-quality data from public resources and organizing cybersecurity question design competitions, and contains 44,823 multiple-choice questions and 3,087 short-answer questions. Furthermore, the dataset uses powerful LLMs for data annotation and building scoring agents to automatically evaluate short-answer questions.




