Agentic Bio-Capabilities Benchmark (ABC-Bench)
收藏资源简介:
ABC-Bench是由SecureBio与Active Site联合创建的智能体生物能力基准测试套件,旨在评估大型语言模型代理在生物安全相关任务中的实际能力。该数据集包含三个核心任务:片段设计、筛选规避和液体处理机器人控制,总计约30个测试样本,数据来源于分子生物学实验模拟与合成生物学设计场景。其创建过程严格遵循七项设计原则,通过算法自动评分和湿实验室验证确保评估的客观性与可重复性。该数据集主要应用于生物安全风险评估与人工智能治理领域,旨在量化AI代理在双用途生物技术任务中的性能,为防范生物恶意滥用提供关键能力测量依据。
ABC-Bench is a benchmark suite for agentic biological capabilities jointly developed by SecureBio and Active Site, designed to evaluate the practical performance of large language model (LLM) agents on biosafety-related tasks. This dataset encompasses three core tasks: fragment design, screening evasion, and liquid-handling robot control, with a total of approximately 30 test samples. The data is derived from molecular biology experiment simulations and synthetic biology design scenarios. Its development strictly adheres to seven design principles, and guarantees the objectivity and reproducibility of the evaluation through automated algorithmic scoring and wet-lab validation. This dataset is primarily utilized in the domains of biosafety risk assessment and AI governance, aiming to quantify the performance of AI agents in dual-use biotechnology tasks and provide a critical capability measurement basis for preventing malicious misuse of biological technologies.




