SecBench
收藏资源简介:
SecBench是一个多维度的基准测试数据集,旨在评估大型语言模型(LLMs)在网络安全领域的表现。该数据集由香港理工大学和腾讯实验室联合创建,包含44,823个多项选择题(MCQs)和3,087个简答题(SAQs),涵盖了中文和英文两种语言,以及知识保留和逻辑推理两个能力水平。数据集的内容来源于公开数据源和网络安全问题设计竞赛,经过GPT-4的自动标注和评分处理。SecBench的应用领域主要是网络安全,旨在解决现有基准测试数据不足、问题形式单一的问题,为LLMs在网络安全领域的性能评估提供全面支持。
SecBench is a multi-dimensional benchmark dataset aimed at evaluating the performance of Large Language Models (LLMs) in the cybersecurity domain. It was jointly developed by The Hong Kong Polytechnic University and Tencent Labs, comprising 44,823 multiple-choice questions (MCQs) and 3,087 short-answer questions (SAQs). The dataset covers two languages (Chinese and English) and two competency levels: knowledge retention and logical reasoning. Its content is sourced from public data resources and cybersecurity problem design competitions, and has undergone automatic annotation and scoring via GPT-4. Primarily applied in the cybersecurity field, SecBench addresses the shortcomings of insufficient existing benchmark datasets and single-form question structures, providing comprehensive support for performance evaluation of LLMs in the cybersecurity domain.

- 1SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity香港理工大学, 腾讯安全科恩实验室, 腾讯朱雀实验室, 腾讯安全平台与部门 · 2024年



