infra-as-code-trajectories
收藏资源简介:
Infra As Code Trajectories 是一个由 Grok 4.6 通过合成数据工厂生成的合成数据集,旨在解决 Terraform 和 Kubernetes 中遗留对象的基础设施即代码(IaC)修复问题。该数据集包含 5208 条原始记录,分布在多个 JSONL 文件中(batch-r01.jsonl 至 batch-r2604.jsonl),总大小约 28699 KB。数据集的发布状态为原始未整理版本,不适用于直接训练,仅供检查和复现。数据集的语言为英语,采用 Apache 2.0 许可证。它被设计为通用代理研究数据集,独立于 Fable 5 集合,且未被标记为 Spikenaut 训练数据。计划中的策展训练版本将在后续审计和导出后发布,当前版本不应作为训练语料库使用。
Infra As Code Trajectories is a synthetic dataset generated by Grok 4.6 via a synthetic data factory, designed to address infrastructure-as-code (IaC) remediation issues for legacy objects in Terraform and Kubernetes. The dataset contains 5208 raw records distributed across multiple JSONL files (batch-r01.jsonl to batch-r2604.jsonl), with a total size of approximately 28699 KB. The release status is the raw, uncurated version, which is not suitable for direct training and is intended only for inspection and reproduction. The dataset language is English, licensed under Apache 2.0. It is designed as a general agent research dataset, independent of the Fable 5 collection, and is not marked as Spikenaut training data. A curated training version is planned for future release after auditing and export; the current version should not be used as a training corpus.
Infra As Code Trajectories 数据集详情
数据集概述
该数据集围绕 Terraform/Kubernetes 遗留对象的基础设施即代码(IaC)修复任务构建,属于合成数据类别,用于代理式工作流研究。
关键信息
- 数据集名称: Infra As Code Trajectories
- 许可证: Apache-2.0
- 语言: 英语
- 数据标签: synthetic-data、agentic-workflows、grok-4.6、provenance、trajectories、terraform、kubernetes、iac
数据来源与生成
- 底层合成数据由 Grok 4.6 通过合成数据工厂代理通道生成
- 工厂标识:
infra-as-code-factory - 源路径:
outputs/raw/2026-08-19-agentic/infra-as-code-factory/
当前发布状态
原始数据已公开,位于 data/raw/ 目录,包含 5208 条记录,分布在 data/raw/batch-r01.jsonl 至 data/raw/batch-r2604.jsonl 文件中(约 28699 KB)。支持性说明文档位于 data/metadata/NOTES-*.md。
重要声明:
- 当前数据为公开原始证据快照,尚未经过整理,不具备训练就绪状态
- 该 Hub 副本并非精选训练导出,工厂源仍是写入目标
- 公开可见性不代表训练就绪性
计划中的精选发布
精选训练数据集的发布仍需等待后续审计和导出流程,此仓库不应被视为训练语料库。
相关资源链接
许可证说明
本公开原始发布采用 Apache License 2.0 许可,该许可授予重用权限,但并不使记录具备训练就绪性或代表真实世界的实际测量数据。




