Full-length and split homologs of human proteins in the gut microbiome
收藏资源简介:
These files were generated as part of the manuscript "Human xenobiotic metabolism proteins have full-length and split homologs in the gut microbiome" (submitted). The .tar file contains .ipc files that are tables of full-length (full_humcover3.ipc) and split homologs (part_humcover3.ipc) of human proteins in the gut microbiome, organized by alignment coverage threshold. For example, the directory `HumanUPR_0.67_src_20000_70` contains results obtained at a 67% alignment coverage threshold for the bacterial protein, and 70% for the human protein. Note that our pipeline collapses full-length alignments to the same UHGP-90 protein family into a single entry per species, with the number of genomes reported in the column nGenomes. Split homologs are not collapsed because genomic context is used to define them, and this context may differ across individual genomes. These files are in Arrow IPC format, which provides compression and fast I/O for large tables. We recommend reading them using pola.rs or the R Arrow package. In particular, because the full-length homolog table is large, you may wish to work with it without loading it into memory, which can be accomplished using scan_ipc in pola.rs or open_dataset in R Arrow. We also provide gzipped .csv format datasets of full-length (pgkb_FH_drugs.csv.gz) and split (pgkb_SH_drugs.csv.gz) homologs, at the default 67% alignment coverage threshold for bacterial and 70% for human proteins, organized by their PharmGKB annotations. For each drug annotated in PharmGKB as being metabolized by a human protein with full-length or split homologs, we provide the human protein(s) responsible, its xenobiotic enzyme class, the bacterial protein homolog(s), length and percent identity of the alignment, and either the specific genome (g, split homologs only) or the number of genomes (nGenomes, full homologs only). Xenobiotic enzyme classes are defined as in Figure 4 of the manuscript, with the additional classes "nucl" (nucleobase-containing metabolic proteins not annotated to any other class), "redox" (oxidoreductases not annotated to any other class), and "other" (all remaining proteins).
本数据集文件源自已投稿的学术论文《人外源性代谢蛋白在肠道菌群中存在全长与截短同源物》(Human xenobiotic metabolism proteins have full-length and split homologs in the gut microbiome)。 该.tar压缩包内含.ipc格式文件,为肠道菌群中人源蛋白的全长同源物(full_humcover3.ipc)与截短同源物(part_humcover3.ipc)的数据表,按序列比对覆盖度阈值(alignment coverage threshold)进行组织。例如,目录`HumanUPR_0.67_src_20000_70`对应的分析结果,其细菌蛋白的比对覆盖度阈值为67%,人源蛋白的比对覆盖度阈值为70%。需注意,本分析流程会将比对至同一UHGP-90蛋白家族的全长比对结果按物种合并为单条记录,物种对应的基因组数量将在nGenomes列中给出。截短同源物则不会被合并,因为其定义依赖基因组上下文(genomic context),而不同个体基因组的上下文可能存在差异。 本数据集文件采用Arrow IPC格式(Arrow IPC format),该格式可对大型数据表实现压缩与快速读写。我们推荐使用pola.rs库或R语言Arrow包进行读取。鉴于全长同源物数据表体积较大,您可选择不将其完全加载至内存中进行分析,这可通过pola.rs中的scan_ipc函数或R Arrow包中的open_dataset函数实现。 我们还提供了gzip压缩的.csv格式数据集,包含默认阈值下的全长同源物(pgkb_FH_drugs.csv.gz)与截短同源物(pgkb_SH_drugs.csv.gz):细菌蛋白比对覆盖度阈值为67%,人源蛋白为70%,并按PharmGKB(PharmGKB)注释进行分类。针对PharmGKB中注释为可被带有全长或截短同源物的人源蛋白代谢的每一种药物,我们提供了如下信息:负责代谢的人源蛋白、其所属的外源性代谢酶类、对应的细菌蛋白同源物、比对长度与序列一致性百分比,以及具体基因组(仅截短同源物,字段为g)或基因组数量(仅全长同源物,字段为nGenomes)。外源性代谢酶类的定义参考论文图4,新增了三类:"nucl"(未注释至其他类别的含核碱基代谢蛋白)、"redox"(未注释至其他类别的氧化还原酶(oxidoreductases))以及"other"(其余所有蛋白)。



