遇见数据集

mindlab-research/CrackedShell-Sec-171

收藏
Hugging Face2026-04-11 更新2026-06-14 收录
官方服务:

资源简介:

--- pretty_name: CrackedShell-Sec-171 license: apache-2.0 tags: - security - benchmark - agents - red-team - skill-bundle size_categories: - 100<n<1K --- # CrackedShell-Sec-171 CrackedShell-Sec-171 is a file-native benchmark corpus for evaluating agent security behavior, runtime safeguards, tool-use controls, and red-team robustness. The dataset ships each benchmark as a standalone bundle so evaluators can preserve the original multi-file layout during ingestion and execution. ## Highlights - 171 benchmark bundles with per-case metadata, labels, and supporting assets. - Multi-file bundle format designed for agent-skill ingestion, runtime-policy testing, and adversarial evaluation pipelines. - Structured annotations exposed through `manifest.json`, `labels.json`, and optional `file_graph.json` files. - Broad scenario coverage across families such as benign (30), unknown (19), data_exfiltration (18), prompt_injection (17), credential_theft (16), agent_hijacking (15). - Severity-style annotations present in the corpus with values including critical (63), high (57), none (33), medium (17), low (1). ## Repository layout ```text . ├── README.md ├── LICENSE └── skill_bundles/ ├── index.json ├── BND-*/ │ ├── manifest.json │ ├── labels.json │ ├── file_graph.json │ └── files/ │ ├── SKILL.md │ ├── docs/ │ ├── runtime/ │ └── scripts/ ``` ## Data format - `skill_bundles/index.json`: corpus-level index for the full benchmark set. - `manifest.json`: case metadata such as family, severity, permissions, entrypoints, and benchmark annotations. - `labels.json`: per-bundle labels for evaluation and filtering. - `file_graph.json`: optional graph representation of relationships between files inside a bundle. - `files/`: benchmark payload content, including instructions, scripts, runtime configs, prompts, and documentation assets. ## Recommended use This corpus is well suited for: - agent security benchmarking - policy and guardrail evaluation - adversarial skill-ingestion tests - tool permission and runtime-control validation - benchmark-driven red-team experiments ## Citation If you use **CrackedShell-Sec-171** in your work, please cite: ```bibtex @misc{shiroyang2026skillssecurity, author = {Shiro Yang and Rio Yang and Andrew Chen and Pony Ma and {Mind Lab}}, title = {OpenClaw Skill Security: Intent-Capability-Behavior Consistency as a New Framework}, year = {2026}, howpublished = {Mind Lab: A Lab for Experiential Intelligence}, note = {https://macaron.im/mindlab/research/openclaw-skill-security-intent-capability-behavior-consistency-framework} } ``` ## License This dataset is released under the Apache License 2.0. You may use, modify, and redistribute the corpus under the terms of the license; see the `LICENSE` file for the full text. Because some benchmark bundles include executable scripts and adversarial test content, they should only be inspected or run in isolated evaluation environments with appropriate safeguards.

--- 数据集名称:CrackedShell-Sec-171 许可证:Apache 2.0 标签: - 安全 - 基准测试 - 智能体 - 红队(red-team) - 技能包 数据集规模:100<n<1K --- # CrackedShell-Sec-171 数据集 CrackedShell-Sec-171 是一款面向文件原生场景的基准语料库,用于评估AI智能体(AI Agent)的安全行为、运行时防护、工具使用控制以及红队(red-team)鲁棒性。本数据集将每个基准测试封装为独立包进行分发,以便评估人员在导入与执行阶段保留原始的多文件布局。 ## 核心特性 - 171个基准包,每个包均附带案例级元数据、标签与配套资源。 - 采用多文件包格式,专为AI智能体技能导入、运行时策略测试以及对抗性评估流水线设计。 - 结构化注释通过`manifest.json`、`labels.json`以及可选的`file_graph.json`文件对外暴露。 - 覆盖广泛的场景类别,包括良性(30)、未知(19)、数据外溢(data_exfiltration)、提示词注入(prompt_injection)、凭据窃取(credential_theft)以及AI智能体劫持(agent_hijacking)等。 - 语料库包含严重程度分级注释,分级取值分别为:严重(63)、高(57)、无(33)、中(17)、低(1)。 ## 仓库布局 text . ├── README.md ├── LICENSE └── skill_bundles/ ├── index.json ├── BND-*/ │ ├── manifest.json │ ├── labels.json │ ├── file_graph.json │ └── files/ │ ├── SKILL.md │ ├── docs/ │ ├── runtime/ │ └── scripts/ ## 数据格式 - `skill_bundles/index.json`:用于完整基准测试集的语料库级索引文件。 - `manifest.json`:存储案例元数据,包括类别、严重程度、权限、入口点以及基准测试注释等信息。 - `labels.json`:存储每个基准包的评估与筛选用标签。 - `file_graph.json`:可选文件,用于表征基准包内部文件间的关联关系图。 - `files/`目录:存储基准测试负载内容,包括指令、脚本、运行时配置文件、提示词以及文档资源等。 ## 推荐应用场景 本语料库适用于以下场景: - AI智能体安全基准测试 - 策略与防护护栏评估 - 对抗性技能导入测试 - 工具权限与运行时控制验证 - 基准测试驱动的红队实验 ## 引用方式 若您在研究工作中使用**CrackedShell-Sec-171**数据集,请引用以下文献: bibtex @misc{shiroyang2026skillssecurity, author = {Shiro Yang and Rio Yang and Andrew Chen and Pony Ma and {Mind Lab}}, title = {OpenClaw Skill Security: Intent-Capability-Behavior Consistency as a New Framework}, year = {2026}, howpublished = {Mind Lab: A Lab for Experiential Intelligence}, note = {https://macaron.im/mindlab/research/openclaw-skill-security-intent-capability-behavior-consistency-framework} } ## 许可证 本数据集采用Apache License 2.0协议发布。您可依据该协议的条款使用、修改并重新分发本语料库;完整协议文本请参见`LICENSE`文件。 由于部分基准包包含可执行脚本与对抗性测试内容,仅可在配备适当防护措施的隔离评估环境中对其进行查看或运行。

提供机构:
mindlab-research
二维码
社区交流群
二维码
科研交流群
商业服务