synthetic-it-support-tickets
收藏资源简介:
Synthetic IT Support Tickets 是一个包含745条合成IT服务管理事件记录的数据集,专为LLM维基生成、知识图谱构建和检索增强生成(RAG)等实验设计。每条记录模拟完整的帮助台或IT运营事件,包括提交的工单信息(如标题、描述、优先级、环境)、带时间戳的故障排除通信记录、结构化的诊断步骤(遵循三步诊断手册)、确定的根本原因以及具体的解决步骤。数据集覆盖14个不同问题族(例如账户锁定、密码重置等),每个问题族包含46至60条记录,平均53.2条。数据通过诊断步骤的结果状态元组(如(fail, pass, pass))来区分同一问题族内的观测变体,共享相同状态元组的工单通常具有相同的根本原因。该数据集支持知识库自动生成、知识图谱信息抽取、基于图谱的RAG系统测试、事件摘要与聚类分析以及合成ITSM工作流程演示等研究与应用。所有数据均为合成生成,包含受控噪声以模拟现实工单的多样性,不应被视为真实生产数据或运营决策依据。
Synthetic IT Support Tickets is a dataset containing 745 synthetic IT service management incident records, designed for experiments such as LLM wiki generation, knowledge graph construction, and retrieval-augmented generation (RAG). Each record simulates a complete help desk or IT operations incident, including submitted ticket information (e.g., title, description, priority, environment), timestamped troubleshooting communication records, structured diagnostic steps (following a three-step diagnostic manual), identified root causes, and specific resolution steps. The dataset covers 14 different problem families (such as account lockout, password reset, etc.), with each family containing 46 to 60 records, averaging 53.2. Data distinguishes observational variants within the same problem family through diagnostic step result state tuples (e.g., (fail, pass, pass)), and tickets sharing the same state tuple often have the same root cause. This dataset supports research and applications like automatic knowledge base generation, knowledge graph information extraction, graph-based RAG system testing, incident summarization and clustering analysis, and synthetic ITSM workflow demonstrations. All data is synthetically generated with controlled noise to simulate the diversity of real tickets and should not be considered as real production data or basis for operational decisions.




