Data from: Distinguishing migration from isolation using genes with intragenic recombination: detecting introgression in the Drosophila simulans species complex
收藏资源简介:
Background: Determining the presence or absence of gene flow between populations is the target of some statistical methods in population genetics. Until recently, these methods either avoided the use of recombining genes, or treated recombination as a nuisance parameter. However, genes with recombination contribute additional information for the detection of gene flow (i.e. through linkage disequilibrium). Methods: We present three summary statistics based on the spatial arrangement of fixed differences, and shared and exclusive polymorphisms that are sensitive to the presence and direction of gene flow. Power and false positive rate for tests based on these statistics are studied by simulation. Results: The application of these tests to populations from the Drosophila simulans species complex yielded results consistent with migration between D. simulans and its two endemic sister species D. mauritiana and D. sechellia, and between populations D. mauritiana on the islands of the Mauritius and Rodrigues. Conclusions: We demonstrate the sensitivity of the developed statistics to the presence and direction of gene flow, and characterize their power as a function of differentiation level and recombination rate. The properties of these statistics make them especially suitable for analyzing high-throughput sequencing data or for their integration within the approximate Bayesian computation framework.
研究背景:判断种群间是否存在基因流,是群体遗传学中若干统计方法的核心研究目标。直至近年,这类方法要么回避使用重组基因,要么将重组视作干扰参数。但携带重组的基因可通过连锁不平衡(linkage disequilibrium)为基因流检测提供额外信息。 研究方法:本文提出三种基于固定差异位点、共享多态位点与专属多态位点空间排布的汇总统计量,这些统计量对基因流的存在与方向具有敏感性。本文通过模拟实验,评估了基于该类统计量的检验方法的检验效能与假阳性率。 研究结果:将这些检验方法应用于拟暗果蝇(Drosophila simulans)物种复合群的种群时,所得结果与拟暗果蝇及其两种特有近缘物种——毛里求斯果蝇(Drosophila mauritiana)和塞舌尔果蝇(Drosophila sechellia)之间的基因流,以及毛里求斯岛与罗德里格斯岛上的毛里求斯果蝇种群间的基因流相符。 研究结论:本文证实了所提出的统计量对基因流的存在与方向具有敏感性,并刻画了其检验效能随分化水平与重组率的变化规律。该类统计量的特性使其特别适用于高通量测序数据分析,或集成至近似贝叶斯计算(approximate Bayesian computation)框架中。



