遇见数据集

Dataset S1: Annotated transcripts in the 1% FST tail

收藏
DataONE2015-07-22 更新2024-06-27 收录
数据链接:
官方服务:

资源简介:

This list contains all transcripts mapped to regions in the 1% tail of FST from the full dataset in either species. Columns represent Ensembl transcript ID, region of high FST where the transcript was annotated, start position of the transcript within the region, end position of the transcript within the region, proportion of the transcript mapped to the region, mean percent identity for parts of the transcript that mapped, Ensembl gene ID, and gene symbol. We annotated genes in the 1% tail of all 20 kb non-overlapping windows by aligning human protein sequences to the blue-eyed black lemur genome. We obtained protein sequences for human genome build hg18 and used TBLASTN version 2.2.22+ (Altschul et al. 1990, 1997), with an e-value threshold of 5 x 10-5 to identify orthologs within the regions of the blue-eyed black lemur reference genome corresponding to the 1% FST tail. We then took the list of all human proteins with hits within candidate regions and performed TBLASTN for these proteins against the entire lemur genome. We retained proteins whose best genome-wide match (containing the lowest e-value or maximum mean percent identity) for any subset of the protein sequence overlapped the candidate region. In cases in which multiple proteins mapped to the same location (>50% protein length overlapping, presumably representing multiple transcripts of the same gene or multiple genes in the same family), we retained the protein with the largest total length spanned by initial TBLASTN hits or the largest mean percent identity.

本列表收录了两个物种种群全数据集内,与FST(遗传分化系数)值最高的1%尾部区域匹配的所有转录本。各列依次为:Ensembl转录本ID(Ensembl transcript ID)、转录本注释所在的高FST区域、转录本在该区域内的起始位置、转录本在该区域内的终止位置、映射至该区域的转录本占比、匹配片段的平均序列一致性百分比、Ensembl基因ID(Ensembl gene ID)以及基因符号。 我们通过将人类蛋白质序列比对至蓝眼黑狐猴基因组,对所有20kb非重叠窗口中处于FST 1%尾部的基因进行了注释。我们获取了人类基因组组装版本hg18的蛋白质序列,并使用TBLASTN 2.2.22+版本(Altschul等人,1990、1997),设置e值阈值为5×10^-5,在蓝眼黑狐猴参考基因组中对应FST值最高1%尾部的区域内识别直系同源基因。 随后,我们提取所有在候选区域内存在比对命中结果的人类蛋白质列表,并将这些蛋白质与整个狐猴基因组进行TBLASTN比对。我们保留了那些其全基因组最优匹配(即e值最低或平均序列一致性百分比最高的匹配结果)的任意蛋白质序列子集与候选区域存在重叠的蛋白质。 当多个蛋白质映射至同一位置(蛋白质序列重叠度超过50%,推测对应同一基因的多个转录本或同一基因家族的多个基因)时,我们保留初始TBLASTN比对结果覆盖总长度最大,或平均序列一致性百分比最高的蛋白质。

创建时间:
2015-07-22
二维码
社区交流群
二维码
科研交流群
商业服务