遇见数据集

Terminal_Safety

收藏
魔搭社区2026-05-27 更新2026-07-15 收录
官方服务:

资源简介:

# Terminal Safety A terminal-agent safety trajectory dataset for training and evaluating guard models. ## Files - `terminal_safety_trajectories_q3_5.jsonl`: filtered ShareGPT-style trajectories with step-level safety labels. ## Dataset Summary - Records: 8,101 - Average turns per trajectory: 14.13 - Turn range: 3 to 164 - Quality filter: 3-model audit, mean quality >= 3.5 ## Outcome Distribution - Safe-Original: 3,522 - Safe-Defended: 1,658 - Unsafe-PartialDefense: 1,096 - Unsafe-Committed: 1,825 ## Risk Categories - R1 Leak: 827 - R2 Unauthorized action/access: 835 - R3 Damage: 775 - R4 Goal/process drift: 723 - R5 Harmful content: 702 - R6 Resource abuse: 717 ## Source Types - User-init: 2,118 - Environment-tool-output: 1,895 - Agent-self-degradation: 566 - Safe original trajectories: 3,522 ## Format Each JSONL row contains a full trajectory in `conversations`, plus metadata such as `outcome`, `source`, `risk`, `suspect_step`, and `risk_step`.

提供机构:
maas
创建时间:
2026-05-22
二维码
社区交流群
二维码
科研交流群
商业服务