cyber-defense-agent-trajectories-sample
收藏资源简介:
该数据集是一个公开的合成网络防御智能体轨迹样本,包含40行数据,来自10个源家族(cd-001至cd-010)。每个家族贡献4行,其中3行标记为训练集,1行标记为验证集。数据行包括智能体可见的任务、工具契约、确定性参考工具轨迹以及最终报告。该样本是从`cyber-defense-public-v1.zip`(SHA-256校验值已给出)中确定性投影得到的,对于每个源家族,选取字典序前三个标记为训练的行和第一个标记为验证的行,且源行保持不变。所有场景均为合成数据,不包含任何真实事件、凭证、恶意软件样本或个人数据,轨迹仅针对声明的合成状态/工具/预言机契约建立行为,不反映隐藏的推理质量或通用网络防御能力。
This dataset is a publicly available sample of synthetic cyber defense agent trajectories, containing 40 rows of data from 10 source families (cd-001 to cd-010). Each family contributes 4 rows, with 3 labeled as training and 1 as validation. The data rows include agent-visible tasks, tool contracts, deterministic reference tool trajectories, and final reports. The sample is deterministically projected from `cyber-defense-public-v1.zip` (SHA-256 checksum provided), selecting the first three lexicographically ordered training rows and the first validation row for each source family, with source rows unchanged. All scenarios are synthetic data, containing no real events, credentials, malware samples, or personal data. The trajectories only establish behavior for the declared synthetic state/tool/oracle contracts and do not reflect hidden reasoning quality or general cyber defense capabilities.
数据集概述
该数据集是 GTDataworks Cyber Defense Rooms 提供的合成网络防御智能体轨迹的公开样本,包含40行数据。
数据规模与划分
- 总共40行,覆盖10个源家族(
cd-001至cd-010) - 其中30行保留源
train划分标签,10行保留源validation划分标签 - 每个家族贡献4行:3行训练数据,1行验证数据
数据内容
- 每行包含智能体可见的任务、工具合约、确定性参考工具轨迹以及最终报告
- 公开的验证行仅用于开发便利,不代表行为验证集;私有评估集不包含在此样本中
数据来源与选择方法
- 样本从
cyber-defense-public-v1.zip确定性投影生成(SHA-256:9c70b42714ec96c20667d5c4368fb941acb85e6e569145afaa5c07c50f21569a) - 对每个源家族,选取按字典序排列的前三行已标记为
train的行,以及第一行已标记为validation的行 - 源行保持不变
- 提供
manifest.json和SHA256SUMS文件作为选择收据和文件哈希
安全性声明
- 所有场景均为合成数据
- 源固定声明不包含任何复制的安全事件、凭据、恶意软件样本或个人数据
- 轨迹仅针对声明的合成状态/工具/预言机合约建立行为,不反映隐藏推理质量或通用网络防御能力
许可信息
- 许可证类型:
gtdataworks-commercial-non-exclusive-license-v1(其他/商业非独家许可)




