Data from: Paralogs are revealed by proportion of heterozygotes and deviations in read ratios in genotyping by sequencing data from natural populations
收藏资源简介:
Whole genome duplications have occurred in the recent ancestors of many plants, fish, and amphibians, resulting in a pervasiveness of paralogous loci and the potential for both disomic and tetrasomic inheritance in the same genome. Paralogs can be difficult to reliably genotype and are often excluded from genotyping-by-sequencing (GBS) analyses; however, removal requires paralogs to be identified which is difficult without a reference genome. We present a method for identifying paralogs in natural populations by combining two properties of duplicated loci: 1) the expected frequency of heterozygotes exceeds that for singleton loci, and 2) within heterozygotes, observed read ratios for each allele in GBS data will deviate from the 1:1 expected for singleton (diploid) loci. These deviations are often not apparent within individuals, particularly when sequence coverage is low; but, we postulated that summing allele reads for each locus over all heterozygous individuals in a population would provide sufficient power to detect deviations at those loci. We identified paralogous loci in three species: Chinook salmon (Oncorhynchus tshawytscha) which retains regions with ongoing residual tetrasomy on eight chromosome arms following a recent whole genome duplication, mountain barberry (Berberis alpina) which has a large proportion of paralogs that arose through an unknown mechanism, and dusky parrotfish (Scarus niger) which has largely re-diploidized following an ancient whole genome duplication. Importantly, this approach only requires the genotype and allele-specific read counts for each individual, information which is readily obtained from most GBS analysis pipelines.
全基因组复制(whole genome duplication)事件曾发生于众多植物、鱼类与两栖动物的近期祖先类群中,由此导致旁系同源基因座(paralogous loci)广泛存在,并使得同一基因组内同时具备二体遗传与四体遗传的潜在可能。旁系同源基因座往往难以实现可靠的基因分型,且常被排除在测序分型(genotyping-by-sequencing,GBS)分析之外;然而,要移除这类基因座,首先需要对其进行识别,而在缺乏参考基因组(reference genome)的情况下,这一过程极具挑战性。我们提出了一种可在自然种群中识别旁系同源基因座的方法,该方法结合了重复基因座的两项核心特征:其一,杂合子(heterozygotes)的预期出现频率高于单拷贝基因座(singleton loci);其二,在GBS数据的杂合个体中,每个等位基因的观测读取比例会偏离二倍体单拷贝基因座所预期的1:1比例。这类偏差通常难以在单个个体中被检测到,尤其是在测序覆盖度(sequence coverage)较低的情况下;但我们推测,将种群中所有杂合个体的某一基因座的等位基因读取数进行汇总,便可获得足够的统计效力以识别这类基因座的偏差。我们在三个物种种群中完成了旁系同源基因座的识别:一是奇努克鲑(Oncorhynchus tshawytscha),其在近期发生全基因组复制后,八条染色体臂上仍保留有持续存在的残留四体性区域;二是山小檗(Berberis alpina),该物种中存在大量由未知机制产生的旁系同源基因座;三是暗绿鹦嘴鱼(Scarus niger),其在古老全基因组复制事件后已基本完成重新二倍体化。值得注意的是,本方法仅需每个个体的基因型数据与等位基因特异性读取计数(allele-specific read counts),而这类信息可从绝大多数GBS分析流程中便捷获取。



