CORAL Benchmarking Datasets
收藏资源简介:
Benchmarking datasets used to train/evaluate Coral - **Positive pairs** were compiled from RNAInter, RPI488, RPI369, RPI2241, and RPI1807. We retained RNAInter interactions tagged as strong experimental evidence; from the RPI datasets we kept only positives. Proteins longer than 1,250 aa and RNAs longer than 6,000 nt were excluded, yielding 24,012 high-confidence positive interactions. - **Splits** come in five folds per strategy: redundant, pairwise at 80/60% and 50/30% RNA/protein identity, component-wise at the same two thresholds (25 folds total). Negative pairs are random RNA-protein re-pairings within the same split, balanced 1:1 with positives.SHA-256: def4ca6d99abcc6177d02f72237f3381b3a2833976a0fe30da04dcfa7dc75234
用于训练与评估Coral的基准测试数据集: - **正样本对(Positive pairs)** 源自RNAInter、RPI488、RPI369、RPI2241及RPI1807数据集。我们保留了RNAInter中标记为强实验证据的相互作用记录;仅从上述RPI系列数据集的阳性样本中筛选。排除长度超过1250个氨基酸的蛋白质与长度超过6000 nt的RNA分子,最终得到24012个高置信度阳性相互作用对。 - **划分方式(Splits)** 每种策略均包含5折交叉验证,具体包括冗余划分、以RNA/蛋白质序列同一性80%/60%和50%/30%为阈值的成对划分,以及在相同两个阈值下的组分划分,总计25折。负样本对为同一划分内随机重新配对的RNA-蛋白质组合,且正负样本数量严格平衡为1:1。 SHA-256: def4ca6d99abcc6177d02f72237f3381b3a2833976a0fe30da04dcfa7dc75234



