multimolecule/bprna-spot-2
收藏资源简介:
bpRNA-spot是一个用于RNA二级结构预测的数据集集合,由SPOT-RNA方法使用。该数据集作为复合仓库发布,包含三个编号组件:bpRNA-spot-0(初始bpRNA分割,来自bpRNA-1m的子集,通过CD-HIT去除序列相似度超过80%的序列,并随机按约8:1:1的比例分为训练、验证和测试集)、bpRNA-spot-1(用于迁移学习的高分辨率PDB RNA分割)和bpRNA-spot-2(仅包含NMR解析的RNA,用于训练后评估)。数据集整体将组件按顺序合并为训练集(TR0 + TR1)、验证集(VL0 + VL1)和测试集(TS0 + TS1 + TS2)。所有二级结构以点括号表示法存储,在转换为点括号表示法之前,移除了可能导致核苷酸与多个伙伴配对的碱基对,并将序列文件中的非A/C/G/U符号标准化为N。数据集包含id、序列、二级结构、结构注释和功能注释等列。
bpRNA-spot is a collection of datasets for RNA secondary structure prediction, which is utilized by the SPOT-RNA method. This dataset is released as a composite repository, comprising three numbered components: bpRNA-spot-0, the initial bpRNA split sourced from bpRNA-1m. Sequences with >80% sequence similarity were filtered out via CD-HIT, and this subset was randomly partitioned into training, validation, and test sets at an approximate ratio of 8:1:1; bpRNA-spot-1, high-resolution PDB-derived RNA splits intended for transfer learning; and bpRNA-spot-2, which exclusively contains RNA structures solved via NMR and is designed for post-training evaluation. For the entire dataset, the components are merged sequentially to form the training set (TR0 + TR1), validation set (VL0 + VL1), and test set (TS0 + TS1 + TS2). All secondary structures are stored in dot-bracket notation. During preprocessing prior to conversion to dot-bracket notation, base pairs that would allow a single nucleotide to pair with multiple partners were removed, and non-A/C/G/U symbols in sequence files were standardized to the character N. The dataset includes columns such as id, sequence, secondary structure, structural annotation, and functional annotation.



