遇见数据集

Hominid Palaeoproteomic Reference Dataset

收藏
Zenodo2023-03-13 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This dataset contains the 'Hominid Palaeoproteomic Reference Dataset'. We used PaleoProPhyler ( https://github.com/johnpatramanis/Proteomic_Pipeline ) to generate a palaeoproteomic reference dataset of protein sequences from ancient and present-day hominids. Using the first two modules of PaleoProPhyler, we translated 195 publicly available whole genomes from extant hominid groups. Details on the processing of the sequences can be found in the supplementary materials of PaleoProPhyler (https://github.com/johnpatramanis/Proteomic_Pipeline/blob/main/GitHub_Tutorial/Supplementary.pdf). We also translated 8 ancient hominin genomes from VCF files, including those of several Neanderthals and one Denisovan. Since the dataset is tailored for palaeoproteomic analyses, we chose to translate proteins that have previously been reported as present in either teeth or bone tissue. We compiled a list of 1,696 proteins from previous works and successfully translated 1,543 of them. For each protein, both the canonical and all alternative protein coding isoforms were translated, leading to a total of 10,058 protein sequences for each individual in the dataset.

本数据集为「人科古蛋白质组参考数据集(Hominid Palaeoproteomic Reference Dataset)」。我们借助PaleoProPhyler工具(https://github.com/johnpatramanis/Proteomic_Pipeline),生成了一套涵盖现生及已灭绝人科物种的古蛋白质组参考蛋白质序列数据集。通过PaleoProPhyler的前两个模块,我们对195份公开可用的现生人科类群全基因组进行了蛋白质序列翻译。序列处理的详细流程可查阅PaleoProPhyler的补充材料(https://github.com/johnpatramanis/Proteomic_Pipeline/blob/main/GitHub_Tutorial/Supplementary.pdf)。此外,我们还从VCF格式文件中翻译了8份古人类基因组,其中包含多份尼安德特人(Neanderthal)基因组及1份丹尼索瓦人(Denisovan)基因组。鉴于本数据集专为古蛋白质组分析打造,我们仅选取了此前已报道可在牙齿或骨骼组织中检出的蛋白质进行翻译。我们从既往研究中整理得到1696种目标蛋白质,最终成功翻译其中1543种。针对每一种目标蛋白质,我们同时翻译了其经典编码亚型及所有可变蛋白编码亚型,最终数据集内每个个体均对应10058条蛋白质序列。

提供机构:
Zenodo
创建时间:
2022-12-06
二维码
社区交流群
二维码
科研交流群
商业服务