oligotox-phase2-dataset
收藏资源简介:
OligoTox Phase 2 Dataset是一个计算生成的、适合AI建模的数据集,用于模拟寡核苷酸相关的肝毒性。该数据集由DBbun LLC发布,作为NIH/NCATS OligoTox开放数据挑战第二阶段的一部分。数据集包含寡核苷酸序列、化学修饰模式、递送平台、剂量/暴露背景、体外或转化肝脏相关检测背景、对照和毒性读数的结构化表格。最终数据集包含1,120条寡核苷酸记录(1,000条生成的非对照寡核苷酸和120条对照)、5,600个检测实例、16,800个重复水平的毒性读数以及127个表格,包括8个核心建模表格(寡核苷酸元数据、聚合化学、位置级修饰、生物物理、剂量、检测、读数和对照)以及每个来源的证据模块和辅助元数据文件。数据集旨在用于开发、训练、基准测试和压力测试寡核苷酸毒性的计算机预测模型。所有值均为计算生成,带有行级来源元数据,区分文献基础的生成值与推断和报告值。
The OligoTox Phase 2 Dataset is a computationally generated, AI-ready dataset designed to model oligonucleotide-related hepatotoxicity. It was released by DBbun LLC as part of the NIH/NCATS OligoTox Open Data Challenge Phase 2. The dataset includes structured tables of oligonucleotide sequences, chemical modification patterns, delivery platforms, dose/exposure contexts, in vitro or translational liver-related assay contexts, controls, and toxicity readouts. The final dataset contains 1,120 oligonucleotide records (1,000 generated non-control oligonucleotides and 120 controls), 5,600 assay instances, 16,800 replicate-level toxicity readouts, and 127 tables, including 8 core modeling tables (oligonucleotide metadata, aggregate chemistry, position-level modifications, biophysical, dose, assay, readout, and control) as well as evidence modules for each source and auxiliary metadata files. The dataset is intended for use in developing, training, benchmarking, and stress-testing in silico predictive models of oligonucleotide toxicity. All values are computationally generated with row-level provenance metadata distinguishing literature-based generated values from inferred and reported values.




