Hominid Palaeoproteomic Reference Dataset
收藏资源简介:
This dataset contains the 'Hominid Palaeoproteomic Reference Dataset'. We used PaleoProPhyler ( https://github.com/johnpatramanis/Proteomic_Pipeline ) to generate a palaeoproteomic reference dataset of protein sequences from ancient and present-day hominids. Using the first two modules of PaleoProPhyler, we translated 195 publicly available whole genomes from extant hominid groups. Details on the processing of the sequences can be found in the supplementary materials of PaleoProPhyler (https://github.com/johnpatramanis/Proteomic_Pipeline/blob/main/GitHub_Tutorial/Supplementary.pdf). We also translated 8 ancient hominin genomes from VCF files, including those of several Neanderthals and one Denisovan. Since the dataset is tailored for palaeoproteomic analyses, we chose to translate proteins that have previously been reported as present in either teeth or bone tissue. We compiled a list of 1,696 proteins from previous works and successfully translated 1,543 of them. For each protein, both the canonical and all alternative protein coding isoforms were translated, leading to a total of 10,058 protein sequences for each individual in the dataset.
本数据集为‘人科古蛋白质组参考数据集(Hominid Palaeoproteomic Reference Dataset)’。我们使用PaleoProPhyler(https://github.com/johnpatramanis/Proteomic_Pipeline)构建了一套涵盖古今人科物种的蛋白质序列古蛋白质组参考数据集。借助PaleoProPhyler的前两个模块,我们对195份公开可得的现存人科类群全基因组进行了蛋白质序列翻译。关于序列处理的详细流程,可查阅PaleoProPhyler的补充材料(https://github.com/johnpatramanis/Proteomic_Pipeline/blob/main/GitHub_Tutorial/Supplementary.pdf)。此外,我们还对8份源自VCF文件的古人类基因组进行了翻译,其中包含多份尼安德特人基因组以及1份丹尼索瓦人基因组。鉴于本数据集专为古蛋白质组分析定制,我们选取了此前已在牙齿或骨组织中被报道存在的蛋白质进行翻译。我们从既往研究中整理得到1696种蛋白质的列表,并成功完成其中1543种蛋白质的翻译。针对每种蛋白质,我们均翻译了其经典编码亚型及所有可变蛋白编码亚型,最终数据集内每个个体对应的蛋白质序列总数达10058条。



