遇见数据集

Different orthology inference algorithms generate similar predicted orthogroups among Brassicaceae species

收藏
DataONE2024-09-11 更新2025-08-23 收录
官方服务:

资源简介:

Premise – Orthology inference is crucial for comparative genomics, and multiple algorithms have been developed to identify putative orthologs for downstream analyses. Despite the abundance of proposed solutions, including publicly available benchmarks, it is difficult to assess which tool to best use for plant species, which commonly have complex genomic histories. Methods – We explored the performance of four orthology inference algorithms – OrthoFinder, SonicParanoid, Broccoli, and OrthNet – on eight Brassicaceae genomes in two groups: one group comprising only diploids and another set comprising the diploids, two mesopolyploids, and one recent hexaploid genome. Results – Orthogroup compositions reflect the species’ ploidy and genomic histories. Additionally, the diploid set had a higher proportion of identical orthogroups. While the diploid+higher ploidy set had a lower proportion of orthogroups with identical compositions, the average degree of similarity between the orthogroups was..., We tested seven variations of four orthology inference algorithms - OrthoFinder, SonicParanoid, Broccoli, and OrthNet. We used publicly available genomes and associated protein files from eight Brassicaceae species and compared the results from these algorithms on two sets of species: one consisting of five diploid species and one consisting of eight species - five of the diploids, two mesopolyploids, and one hexaploid. We examined the similarities and differences in the results to understand the performance of these algorithms in inferring orthologs and orthogroups with and without genomically complex species., , # **Different orthology inference algorithms generate similar predicted orthogroups among Brassicaceae species** [https://doi.org/10.5061/dryad.8sf7m0cw8](https://doi.org/10.5061/dryad.8sf7m0cw8) # Description of the data and file structure Primary data, scripts, and outputs are found on GitHub: [https://github.com/itliao/OrthologyComparison](https://github.com/itliao/OrthologyComparison) The first page outlines the types of files in each of the directories, and each directory has a separate README that walks through the steps of the analyses, important scripts, inputs for the scripts, and outputs generated from the scripts. Some of the input/output files are too large to host on GitHub, and thus are found in this repository: ## Gene\_Composition\_Comparison\_Orthogroups - OUTPUTS The following files are outputs from making comparisons between the orthogroup gene compositions from two different orthology inference algorithms. The \"diploid set\" refers to orthogroup inferences made...

# 研究背景 直系同源基因推断(orthology inference)是比较基因组学的核心任务,目前已开发出多种算法用于识别推定直系同源基因(putative orthologs)以支撑下游分析。尽管已有大量相关解决方案(包括公开可用的基准测试集),但针对通常拥有复杂基因组演化历史的植物物种,如何选择最优的分析工具仍是一大难题。 # 研究方法 本研究针对8个十字花科(Brassicaceae)基因组,将其分为两组,测试了4种直系同源基因推断算法——OrthoFinder、SonicParanoid、Broccoli及OrthNet——的性能:第一组仅包含二倍体(diploid)物种,第二组则包含二倍体物种、2个中多倍体(mesopolyploid)物种以及1个新近形成的六倍体(hexaploid)物种。 本研究同时测试了上述4种算法的7种变体。我们使用了公开获取的8个十字花科物种的基因组及对应蛋白质组文件,并针对两组物种分别对比各算法的分析结果:第一组为5个二倍体物种,第二组则包含前述5个二倍体物种、2个中多倍体物种及1个六倍体物种。通过分析结果间的异同,我们旨在明确这些算法在包含/不包含基因组复杂物种的场景下,推断直系同源基因及直系同源基因簇(orthogroup)的性能表现。 # 研究结果 直系同源基因簇的组成能够反映物种的倍性水平及基因组演化历史。此外,仅包含二倍体物种的组别中,拥有完全一致基因组成的直系同源基因簇占比更高。而包含二倍体与高倍性物种的组别中,基因组成完全一致的直系同源基因簇占比更低,但直系同源基因簇间的平均相似程度为…… 我们通过对比分析结果的异同,以探究这些算法在推断直系同源基因及直系同源基因簇时的性能,尤其是在包含或不包含基因组复杂物种的场景下。 # **不同直系同源基因推断算法在十字花科物种间可生成相似的预测直系同源基因簇** [https://doi.org/10.5061/dryad.8sf7m0cw8](https://doi.org/10.5061/dryad.8sf7m0cw8) # 数据与文件结构说明 原始数据、分析脚本及分析结果均已上传至GitHub:[https://github.com/itliao/OrthologyComparison](https://github.com/itliao/OrthologyComparison)。首页概述了各子目录内的文件类型,每个子目录均配有独立的README文档,用于说明分析流程、核心脚本、脚本输入文件及生成的输出文件。部分输入/输出文件体积过大,无法上传至GitHub,因此存放于本数据集仓库中。 ## 直系同源基因簇基因组成对比——输出文件 以下文件为针对两种不同直系同源基因推断算法所得到的直系同源基因簇基因组成进行对比后生成的输出文件。其中"二倍体组别"指基于二倍体物种所得到的直系同源基因簇推断结果……

创建时间:
2025-08-04
二维码
社区交流群
二维码
科研交流群
商业服务