遇见数据集

CURIA cross-species ncRNA correspondence predictions across 19 mammalian genomes

收藏
Zenodo2026-07-15 更新2026-08-01 收录
官方服务:

资源简介:

This dataset contains the per-species output directories used for the manuscript“Cross-species ncRNA annotation using synteny-constrained embedding similarity.” CURIA combines genome-alignment-chain-guided candidate-region projection withRNA foundation-model embeddings to identify localized candidate correspondencesbetween human ncRNA loci and 19 mammalian query genomes. The archive includes thefinal result snapshot used for the reported analyses, including candidate-regionclassifications, reference and query island coordinates, accepted localizedmatches, compact-ncRNA predictions, union-transcript mappings, and associatedrun metadata. The reference assembly is human GRCh38/hg38. Query assemblies include mm39,rn7, bosTau9, felCat9, canFam5, equCab3, susScr11, rheMac10, dasNov3, monDom5,eriEur2, and eight additional mammalian assemblies represented by their assemblyidentifiers in the archive. Each `hg38_vs_<assembly>` directory contains the complete output for onehuman-to-query comparison. BED coordinates use the 0-based, half-open convention.Lower embedding–Smith–Waterman distance values indicate stronger localizedembedding similarity. These results represent candidate cross-species ncRNA correspondences and shouldnot be interpreted independently as definitive evidence of orthology, conservedsecondary structure, or conserved molecular function. The corresponding software, documentation, and download helper are available inthe CURIA repository. This Zenodo record provides the immutable result snapshotused for the preprint.

提供机构:
Zenodo
创建时间:
2026-07-15
二维码
社区交流群
二维码
科研交流群
商业服务