Hominid Palaeoproteomic Reference Dataset
收藏资源简介:
This dataset contains the 'Hominid Palaeoproteomic Reference Dataset'. We used PaleoProPhyler ( https://github.com/johnpatramanis/Proteomic_Pipeline ) to generate a palaeoproteomic reference dataset of protein sequences from ancient and present-day hominids. Using the first two modules of PaleoProPhyler, we translated 195 publicly available whole genomes from extant hominid groups. Details on the processing of the sequences can be found in the supplementary materials of PaleoProPhyler ( https://github.com/johnpatramanis/Proteomic_Pipeline/blob/main/GitHub_Tutorial/Supplementary.pdf ). We also translated 8 ancient hominin genomes from VCF files, including those of several Neanderthals and one Denisovan. Since the dataset is tailored for palaeoproteomic analyses, we chose to translate proteins that have previously been reported as present in either teeth or bone tissue. We compiled a list of 1,696 proteins from previous works and successfully translated 1,543 of them. For each protein, both the canonical and all alternative protein coding isoforms were translated, leading to a total of 10,058 protein sequences for each individual in the dataset. Content: The zipped file contains 4 files, two fasta files as well as two additional folders: - PalaeoProPhyler_Publication_Data_for_Tree.fa contains all of the sequences used to generate the phylogenetic tree presented at PalaeoProPhylers manuscript - ALL_PROT_REFERENCE.fa contains all of the sequences generated as part of the Hominid Palaeoproteomic Reference Dataset described above - PER_PROTEIN is a folder containing one fasta file for each protein within the Hominid Palaeoproteomic Reference Dataset, each protein fasta file has the sequences of all individuals for that particular protein - PER_SAMPLE is a folder containing one fasta file for each sample/individual within the Hominid Palaeoproteomic Reference Dataset, each sample fasta file has the sequences of all proteins for that particular sample. Important Note: The dataset is still not <em>fully </em>generated yet. However around 170 modern samples are complete and ready to be used in their appropriate folder (~/PER_SAMPLE /).
本数据集为**人科古蛋白质组参考数据集(Hominid Palaeoproteomic Reference Dataset)**。我们使用PaleoProPhyler(https://github.com/johnpatramanis/Proteomic_Pipeline)构建了一套涵盖古今人科物种的蛋白质序列古蛋白质组参考数据集。借助PaleoProPhyler的前两个分析模块,我们对195个公开可获取的现存人科类群全基因组进行了翻译。关于序列处理的详细细节,可查阅PaleoProPhyler的补充材料(https://github.com/johnpatramanis/Proteomic_Pipeline/blob/main/GitHub_Tutorial/Supplementary.pdf)。 我们还对来自VCF格式文件的8个古人类基因组进行了翻译,其中包含多例尼安德特人(Neanderthal)及1例丹尼索瓦人(Denisovan)。鉴于本数据集专为古蛋白质组分析定制,我们仅翻译了此前已在牙齿或骨骼组织中被报道存在的蛋白质。我们从已有研究中整理得到1696种蛋白质的列表,并成功翻译其中1543种。针对每种蛋白质,我们均翻译了其经典剪接体及所有可变剪接编码亚型,因此数据集中每个个体对应总计10058条蛋白质序列。 数据集内容说明:本压缩包包含4个文件与2个附加文件夹: 1. `PalaeoProPhyler_Publication_Data_for_Tree.fa`:包含用于构建PaleoProPhyler论文中系统发育树的全部序列; 2. `ALL_PROT_REFERENCE.fa`:包含本数据集所有已生成的蛋白质序列; 3. `PER_PROTEIN`文件夹:按蛋白质分类存储的FASTA文件,每个文件对应一种蛋白质,包含该蛋白质在所有个体中的序列; 4. `PER_SAMPLE`文件夹:按样本/个体分类存储的FASTA文件,每个文件对应一个样本,包含该个体的所有蛋白质序列。 重要说明:本数据集尚未完全构建完成,但约170个现代样本已完成构建,可在对应文件夹(~/PER_SAMPLE/)中直接使用。



