遇见数据集

Hominid Palaeoproteomic Reference Dataset

收藏
Zenodo2023-03-13 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This dataset contains the 'Hominid Palaeoproteomic Reference Dataset'. We used PaleoProPhyler ( https://github.com/johnpatramanis/Proteomic_Pipeline ) to generate a palaeoproteomic reference dataset of protein sequences from ancient and present-day hominids. Using the first two modules of PaleoProPhyler, we translated 195 publicly available whole genomes from extant hominid groups. Details on the processing of the sequences can be found in the supplementary materials of PaleoProPhyler (https://github.com/johnpatramanis/Proteomic_Pipeline/blob/main/GitHub_Tutorial/Supplementary.pdf). We also translated 8 ancient hominin genomes from VCF files, including those of several Neanderthals and one Denisovan. Since the dataset is tailored for palaeoproteomic analyses, we chose to translate proteins that have previously been reported as present in either teeth or bone tissue. We compiled a list of 1,696 proteins from previous works and successfully translated 1,543 of them. For each protein, both the canonical and all alternative protein coding isoforms were translated, leading to a total of 10,058 protein sequences for each individual in the dataset.

本数据集为古人类蛋白质组参考数据集(Hominid Palaeoproteomic Reference Dataset)。我们借助PaleoProPhyler(https://github.com/johnpatramanis/Proteomic_Pipeline)构建了一套覆盖古今人类的古蛋白质组参考序列数据集。通过PaleoProPhyler的前两个模块,我们对195份公开可得的现生人类类群全基因组进行了翻译。序列处理的详细流程可参见PaleoProPhyler的补充材料(https://github.com/johnpatramanis/Proteomic_Pipeline/blob/main/GitHub_Tutorial/Supplementary.pdf)。我们还从VCF文件中翻译了8份古人类基因组,其中包含多份尼安德特人基因组与1份丹尼索瓦人基因组。鉴于本数据集专为古蛋白质组分析定制,我们选取了此前文献中报道过的、可在牙齿或骨骼组织中检出的蛋白质开展翻译工作。我们从既往研究中整理得到1696种目标蛋白,最终成功翻译其中1543种。针对每种蛋白,我们同时翻译了其经典亚型及所有可变蛋白编码亚型,最终数据集内每个个体对应总计10058条蛋白质序列。

提供机构:
Zenodo
创建时间:
2022-12-06
二维码
社区交流群
二维码
科研交流群
商业服务