OpsEval Dataset
收藏资源简介:
OpsEval数据集代表了在评估IT运营中的人工智能(AIOps)方面的一项开创性工作,专注于大型语言模型(LLMs)在此领域的应用。在IT运营越来越依赖AI技术进行自动化和效率提升的时代,理解LLMs在运营任务中的性能变得至关重要。OpsEval提供了一个全面的任务导向基准,专门设计用于评估LLMs在各种关键IT运营场景中的表现。
The OpsEval dataset represents a pioneering effort in evaluating Artificial Intelligence for IT Operations (AIOps), with a focus on the application of Large Language Models (LLMs) in this domain. In an era where IT operations increasingly rely on AI technologies for automation and efficiency enhancement, understanding the performance of LLMs in operational tasks has become crucial. OpsEval provides a comprehensive task-oriented benchmark specifically designed to assess the performance of LLMs across various critical IT operational scenarios.
OpsEval Dataset 概述
数据集简介
OpsEval 数据集是一项针对人工智能IT运维(AIOps)评估的开创性工作,专注于大型语言模型(LLMs)在该领域的应用。该数据集提供了一个全面的任务导向基准,用于评估LLMs在各种关键IT运维场景中的性能。
数据集亮点
- 全面评估:包含7184个多选题和1736个问答格式,支持中英文,是AIOps领域中最全面的基准之一。
- 任务导向设计:专门设计用于评估LLMs在不同关键场景和能力水平上的熟练度。
- 专家审核:数十名领域专家手动审核问题,确保评估的可靠性。
- 开源与动态排行榜:已开源20%的测试问答,便于研究人员进行初步评估。实时更新的在线排行榜记录了新兴LLMs的性能。
数据集结构
/dev/:用于少样本上下文学习的示例。/test/:OpsEval的测试集。
数据集信息
| 数据集名称 | 开源大小 |
|---|---|
| 有线网络 | 1563 |
| Oracle数据库 | 395 |
| 5G通信 | 349 |
| 日志分析 | 310 |
引用信息
当在研究中引用OpsEval数据集时,请使用以下引用格式:
@misc{liu2024opseval, title={OpsEval: A Comprehensive IT Operations Benchmark Suite for Large Language Models}, author={Yuhe Liu and Changhua Pei and Longlong Xu and Bohan Chen and Mingze Sun and Zhirui Zhang and Yongqian Sun and Shenglin Zhang and Kun Wang and Haiming Zhang and Jianhui Li and Gaogang Xie and Xidao Wen and Xiaohui Nie and Minghua Ma and Dan Pei}, year={2024}, eprint={2310.07637}, archivePrefix={arXiv}, primaryClass={cs.AI} }




