multimolecule/bprna-spot-0
收藏资源简介:
bpRNA-spot是一个用于RNA二级结构预测的数据集集合,由SPOT-RNA使用。该数据集作为一个复合存储库发布,包含三个编号的组件存储库:bpRNA-spot-0(初始bpRNA分割,包括TR0、VL0和TS0)、bpRNA-spot-1(PDB迁移学习分割,包括TR1、VL1和TS1)和bpRNA-spot-2(仅NMR评估分割,包括TS2)。这些组件按顺序拼接成训练集(TR0 + TR1)、验证集(VL0 + VL1)和测试集(TS0 + TS1 + TS2)。TR0/VL0/TS0分割是bpRNA-1m的子集,通过CD-HIT(CD-HIT-EST)去除序列相似性超过80%的序列,并将剩余序列按约8:1:1的比例随机分割为训练、验证和测试集。TR1/VL1/TS1分割包含用于迁移学习的高分辨率PDB RNA。TS2分割包含39个通过NMR解析的RNA,用于后训练评估。所有二级结构均以点括号符号存储。在序列/标签分割中,在转换为点括号符号之前,会移除可能导致核苷酸与多个伙伴配对的碱基对。序列文件中的非A/C/G/U符号被归一化为N。数据集包含以下列:id(序列标识符)、sequence(RNA序列)、secondary_structure(点括号符号表示的二级结构,可能使用超出()的括号层级表示假结)、structural_annotation(基于存储的点括号结构生成的bpRNA样式结构注释)和functional_annotation(基于存储的点括号结构生成的bpRNA样式功能注释)。该数据集主要用于文本生成和掩码语言建模任务,支持语言建模和掩码语言建模。
bpRNA-spot is a collection of datasets for RNA secondary structure prediction used by SPOT-RNA. Released as a composite repository, it contains three numbered component repositories: bpRNA-spot-0 (the initial bpRNA split including TR0, VL0 and TS0), bpRNA-spot-1 (the PDB transfer learning split including TR1, VL1 and TS1) and bpRNA-spot-2 (the NMR-only evaluation split including TS2). These components are concatenated to form the training set (TR0 + TR1), validation set (VL0 + VL1) and test set (TS0 + TS1 + TS2). The TR0/VL0/TS0 split is a subset of bpRNA-1m, where sequences with over 80% sequence similarity are removed using CD-HIT (CD-HIT-EST), and the remaining sequences are randomly split into training, validation and test sets at an approximate ratio of 8:1:1. The TR1/VL1/TS1 split contains high-resolution PDB RNAs for transfer learning. The TS2 split includes 39 NMR-resolved RNAs for post-training evaluation. All secondary structures are stored in dot-bracket notation. During sequence/label splitting, base pairs that could cause a nucleotide to pair with multiple partners are removed prior to conversion to dot-bracket notation. Non-A/C/G/U symbols in sequence files are normalized to N. The dataset includes the following columns: id (sequence identifier), sequence (RNA sequence), secondary_structure (secondary structure represented by dot-bracket notation, where bracket layers beyond standard parentheses may be used to denote pseudoknots), structural_annotation (bpRNA-style structural annotations generated from the stored dot-bracket structure) and functional_annotation (bpRNA-style functional annotations generated from the stored dot-bracket structure). This dataset is primarily used for text generation and masked language modeling tasks, supporting both language modeling and masked language modeling.



