CPSDBench
收藏资源简介:
CPSDBench是一个专为评估中文公共安全领域大型语言模型性能而设计的数据集。由中国人民公安大学等机构的研究人员开发,该数据集整合了来自真实世界场景的公共安全相关数据,支持对大型语言模型在文本分类、信息提取、问答和文本生成等任务上的全面评估。数据集包含4000个样本,旨在通过模拟真实世界中遇到的复杂情况,更详细和全面地探索大型语言模型解析和生成与公共安全相关数据的能力。此外,该研究还引入了一套创新的评估指标,旨在更精确地量化大型语言模型在执行公共安全相关任务时的效能。
CPSDBench is a dataset specifically designed for evaluating the performance of large language models (LLMs) in the Chinese public safety domain. Developed by researchers from institutions including the People's Public Security University of China, this dataset integrates public safety-related data sourced from real-world scenarios, enabling comprehensive assessments of LLMs across multiple tasks such as text classification, information extraction, question answering, and text generation. The dataset comprises 4,000 samples, aiming to conduct a more detailed and comprehensive exploration of LLMs' abilities to parse and generate public safety-related data by simulating complex real-world scenarios. Furthermore, this study proposes an innovative set of evaluation metrics, which are intended to more precisely quantify the efficacy of LLMs when executing public safety-related tasks.

- 1CPSDBench: A Large Language Model Evaluation Benchmark and Baseline for Chinese Public Security Domain中国人民公安大学 · 2024年



