Synthetic Incident Response Log Creation with Large Language Models
收藏资源简介:
This dataset presents a gold standard collection of synthetic cybersecurity incident response activities and communication logs datasets, created to address the critical shortage of open source publicly available incident response datasets, including synthetic data detailing incident responder activity. The data was generated and refined through a large language model ChatGPT augmentation driven process, guided by a novel error taxonomy and evaluation framework developed in this research. A file containing all prompts to create each dataset has also been contributed. The dataset includes: Incident response process activities. Analyst communication logs simulating team interactions during incident handling. Structured fields such as case identifiers, timestamps, state transitions, priorities, and observables. Refined annotations validated by cybersecurity professionals to establish a reproducible ground truth. The dataset is designed to be used as: A benchmark resource for evaluating large language model outputs in the context of incident response datasets. A research enabling resource to advance cyber defense, by enabling systematic studies of data quality, error patterns, and human-Artificial Intelligence collaboration in incident response. This work contributes to filling a major data void in cybersecurity research by providing a reproducible, validated, and open source publicly available gold standard collection of synthetic cybersecurity incident response activities and communication logs datasets.



