遇见数据集

An integrative workflow for lncRNA orthology detection and its application to 13 evolutionarily diverse species - Example with 3 species

收藏
Zenodo2026-03-16 更新2026-05-29 收录
官方服务:

资源简介:

This dataset contains the intermediate and final files generated by the workflow described in the manuscript “An integrative workflow for lncRNA orthology detection and its application to 13 evolutionarily diverse species”. The objective of this workflow is to identify putative orthologous long non-coding RNAs (lncRNAs) across vertebrate species by combining complementary genomic conservation signals. The approach integrates synteny-based inference and sequence conservation derived from multi-species genome alignments, enabling both pairwise and multi-species analyses of lncRNA conservation. In the paper, the analyses were performed on 13 vertebrate species including model organisms (human, mouse, zebrafish) and several domestic or farm animals (dog, cow, horse, pig, goat, chicken, duck, turkey, zebra finch and ostrich) but here, for memory purposes, only three specieis (human, mouse, dg) have been considered. The workflow relies on Ensembl genome annotations and one-to-one orthologous relationships between protein-coding genes (PCGs). Three complementary strategies were implemented to infer putative lncRNA orthology relationships: Orthology by conserved genomic context (synteny) using two flanking orthologous PCGs. Orthology by synteny and lncRNA–PCG orientation, based on the FEELnc classification relative to the nearest PCG. Orthology by genome alignment, using conserved genomic segments obtained from Ensembl Compara multi-species alignments generated by the Mercator–Pecan algorithm. The files provided here correspond to the different stages of the workflow and include processed genomic annotations, intermediate tables produced by each orthology inference method, and integrated summaries combining results across methods and species. These tables report the inferred relationships between lncRNAs (one-to-one, one-to-many, many-to-many) as well as the level of support provided by the different methods. In addition to orthology inference, the dataset includes downstream analyses used to explore the functional conservation of candidate lncRNAs. These analyses include detection of conserved short sequence motifs using LncLOOM and cross-species expression analyses based on RNA-seq datasets from human and chicken across a panel of homologous tissues. Together, this dataset provides a comprehensive resource to explore lncRNA conservation across vertebrates and to reproduce the analyses described in the manuscript. The workflow and associated scripts are available in the accompanying GitHub repository, enabling users to reproduce the analyses or apply the approach to other species supported by Ensembl annotations.

提供机构:
Zenodo
创建时间:
2026-03-16
二维码
社区交流群
二维码
科研交流群
商业服务