jang1563/bio-constitution-rules
收藏资源简介:
Bio Constitution Rules — Synthetic Training Corpus是一个包含1,063条记录的标记训练语料库,用于生物双重用途内容分类。每条记录都带有双轨标签:生物特定的真实标签和通用的CBRN基线标签。39.3%的记录在这两个系统之间存在差异,这些是最有价值的训练示例,捕捉了生物特定规则独特编码的区别。数据集包含多个领域,如病毒学、毒理学、合成生物学等,每个领域都有特定的规则和记录数量。数据集还详细描述了记录的架构、标签类型、差异类别以及生成方法。
Bio Constitution Rules — Synthetic Training Corpus is a 1,063-record labeled training corpus for biological dual-use content classification, generated from 30 bio-domain constitutional rules. Each record carries dual-track labels: a bio-specific ground truth label and a generic CBRN baseline label. 39.3% of records diverge between the two systems — these are the highest-value training examples, capturing the distinctions that bio-specific rules uniquely encode. The dataset covers multiple domains such as virology, toxicology, synthetic biology, etc., each with specific rules and record counts. It also details the record schema, label types, divergence categories, and generation methodology.




