BioDisclose
收藏资源简介:
BioDisclose是由多所高校研究团队联合构建的面向生物医学安全领域的对抗性诱导基准数据集,旨在评估大语言模型在复杂语境下对敏感生物医学知识的披露行为。该数据集包含480条精心设计的单轮提示,涵盖病原生物学、人类基因编辑、合成生物学等六大风险领域,并采用学术、历史、角色扮演和分步分解四种诱导策略生成。数据通过专家撰写场景、结构化表示与多轮质量控制流程构建,确保了科学合理性和风险代表性。该数据集主要应用于评估大语言模型在生物医学研究场景中的安全边界,为解决双用途知识泄露风险提供细粒度的技术披露度量标准。
BioDisclose is an adversarial prompting benchmark dataset for the field of biomedical security, jointly constructed by research teams from multiple universities. It aims to evaluate the disclosure behavior of Large Language Models (LLMs) regarding sensitive biomedical knowledge in complex contexts. The dataset contains 480 well-designed single-turn prompts, covering six high-risk domains including pathogenic biology, human genome editing, synthetic biology and other related fields. It is generated using four prompting strategies: academic inquiry, historical context, role-playing and step-by-step decomposition. The dataset is built through expert-authored scenarios, structured representation and a multi-round quality control workflow, ensuring scientific plausibility and risk representativeness. It is mainly applied to evaluate the safety boundaries of LLMs in biomedical research scenarios, and provides a fine-grained technical disclosure metric for mitigating the risk of dual-use knowledge leakage.
数据集摘要:BioDisclose
- 数据集名称:BioDisclose: An Actionability-Aware Benchmark for Biomedical Safety under Adversarial Elicitation
- 核心目的:用于衡量大型语言模型(LLM)在对抗性诱导下,披露生物医学双用途知识的程度。
- 数据集规模与构成:
- 包含 480 个提示(prompts)。
- 这些提示源自 24 个由专家编写的场景。
- 覆盖 6 个生物医学风险领域。
- 涵盖 4 类诱导策略家族:学术、历史、角色扮演和分解提示。
- 评估方法:
- 采用四级评分量表对模型回答进行分级,从“拒绝”到“可执行的披露”。
- 区分高级讨论、技术性具体内容和可执行内容。
- 识别“先拒绝后泄露”的行为。
- 主要发现:
- 在五个已部署的LLM系统中,详细或更高级别的披露率差异显著,范围从 9.2% 到 64.0%。
- 学术框架是最有效的诱导方式,平均成功率为 43.2%。
- 实验室安全场景在所有领域中的披露率最高,达到 51.5%。
- 结果表明,当前的安全保障措施在高风险的科学环境中仍然不一致。
- 论文作者:Yinuo Zhu, He Liu, Boyuan Gu
- 提交日期:2026年7月28日
- arXiv编号:
arXiv:2607.25700v1 - 学科分类:Human-Computer Interaction (cs.HC)




