ibm-research/synthetic-conversations-traces
收藏资源简介:
该数据集包含720个合成的OpenTelemetry(OTel)跟踪记录,模拟用户与大型语言模型(LLM)助手之间的多轮聊天对话,覆盖12个日常建议主题,每个跟踪记录包含30到50轮对话。对话是使用llama-3-3-70b-instruct模型生成的:脚本化的用户消息驱动每一轮,同时模型实时产生助手响应,每轮对话被记录为一个OTel跨度,包括完整的累积消息历史。该数据集专为测试和开发可观察性工具而设计,这些工具遵循gen_ai.*语义约定,用于处理生成式人工智能的遥测数据。
This dataset contains 720 synthetic OpenTelemetry (OTel) trace records that simulate multi-turn chat conversations between users and Large Language Model (LLM) assistants, spanning 12 daily advice-related topics. Each trace record includes 30 to 50 conversation turns. The conversations were generated using the llama-3-3-70b-instruct model: scripted user messages drive each conversation turn, while the model generates assistant responses in real time. Each conversation turn is recorded as an OTel span, which contains the full accumulated message history. This dataset is specifically designed for testing and developing observability tools that follow the gen_ai.* semantic conventions for processing generative AI telemetry data.




