Microbiomehd: The Human Gut Microbiome In Health And Disease
收藏资源简介:
<strong>Overview</strong> MicrobiomeHD is a standardized database of human gut microbiome studies in health and disease. This database includes publicly available 16S data from published case-control studies and their associated patient metadata. Raw sequencing data for each study was downloaded and processed through a standardized pipeline. To be included in MicrobiomeHD, datasets have: publicly available raw sequencing data (fastq or fasta) publicly available metadata with at least case and control labels for each patient at least 15 case patients Currently, MicrobiomeHD is focused on stool samples. Additional samples may be included in certain datasets, as indicated in the metadata. <strong>Files</strong> Additional information about the datasets included in this MicrobiomeHD release are in the MicrobiomeHD github repo https://github.com/cduvallet/microbiomeHD, in the file <em>db/dataset_info.yaml</em>. Top-level identifiers correspond to the dataset IDs used in Duvallet et al. 2017. Sample sizes in the yaml file are those that were described in the papers, and may not exactly reflect the actual data (due to missing/extra data, samples which didn't pass quality control, etc). Each dataset was downloaded and processed through a standardized pipeline. The raw processing results are available in the *.tar.gz files here. Each file has the same directory structure and files, as described in the pipeline documentation: http://amplicon-sequencing-pipeline.readthedocs.io/en/latest/output.html. Specific files of interest include: <strong>summary_file.txt</strong>: this file contains a summary of all parameters used to process the data <strong>datasetID.metadata.txt</strong>: the metadata associated with the samples. Note that some samples in the metadata may not have sequencing data, and vice versa. <strong>RDP/datasetID.otu_table.100.denovo.rdp_assigned</strong>: the 100% OTU tables with Latin taxonomic names assigned using the RDP classifier. <strong>datasetID.otu_seqs.100.fasta</strong>: representative sequences for each OTU in the 100% OTU table. OTU labels in the OTU table end with d__denovoID - these denovoIDs correspond to the sequences in this file. Processing The raw data was acquired as described in the supplementary materials of Duvallet et al.'s "Meta analysis of microbiome studies identifies shared and disease-specific patterns". Raw sequencing data was processed with the Alm lab's in-house 16S processing pipeline: https://github.com/thomasgurry/amplicon_sequencing_pipeline Pipeline documentation is available at: http://amplicon-sequencing-pipeline.readthedocs.io/ Metadata was extracted from the original papers and/or data sources, and formatted manually. <strong>Contributing</strong> MicrobiomeHD is a resource that can be used to extract disease-specific microbiome signals in individual case-control studies. Many microbes respond non-specifically to health and disease, and the majority of bacterial associations within individual studies overlap with this "core" response. Researchers should cross-check their results with the data presented here to ensure that their identified microbial associations are specific to their disease under study. We provide an updated list of "core" microbes here, as well as the raw OTU tables for anyone who wishes to reproduce and adapt this analysis to their study question. If you would like to include your case-control dataset in MicrobiomeHD, please email duvallet[at]mit.edu. For us to process your data through our standard pipeline, you will need to provide the following files and information about your data: raw sequencing data in fastq or fasta format (preferably fastq) information about which processing steps will be required (e.g. removing primers or barcodes, merging paired-end reads, etc) sample IDs associated with the sequencing data (either mapped to barcodes still in the sequences, or to each de-multiplexed sequencing file) case/control metadata of each sample other relevant metadata (e.g. sampling site, if not all samples are stool; sampling time point, if multiple samples per patient were taken; etc) By using MicrobiomeHD in your own analyses, you agree to contribute your dataset to this database and to make your raw sequencing data (i.e. fastq files) publicly available. <strong>Citing MicrobiomeHD</strong> The MicrobiomeHD database and original publications for each of these datasets are described in Duvallet et al. (2017): http://biorxiv.org/content/early/2017/05/08/134031 If you use any of these datasets in your analysis, please cite both MicrobiomeHD (Duvallet et al. (2017)) and the original publication for each dataset that you use. The code used to process and analyze this data in Duvallet et al. (2017) is available on github: https://github.com/cduvallet/microbiomeHD <strong>Files</strong> <em>Core genera</em> <strong>file-S3.core_genera.txt</strong>: Supplemental Table 3 from Duvallet et al. (2017), listing the core health- and disease-associated microbes. <em>Datasets</em> Note that MicrobiomeHD contains all 28 datasets from Duvallet et al. (2017), as well as additional datasets which did not meet the inclusion criteria for the meta-analysis presented in the paper. Additional information about the datasets included in this MicrobiomeHD release are in the original publications and the MicrobiomeHD github repo https://github.com/cduvallet/microbiomeHD, in the file <em>db/dataset_info.yaml</em>. The sample sizes listed here reflect what was reported in the original publications. Some may have discrepancies between what is reported and what is in the actual data due to missing data, quality issues, barcode mismatches, etc. <strong>asd_son_results.tar.gz</strong> (<em>asd_son</em>): NT: 44, ASD: 59 http://dx.doi.org/10.1371/journal.pone.0137725 <strong>autism_kb_results.tar.gz</strong> (<em>asd_kang</em>): H: 20, ASD: 20 http://dx.doi.org/10.1371/journal.pone.0068322 <strong>cdi_schubert_results.tar.gz</strong> (<em>noncdi_schubert</em>): H: 155, nonCDI: 89, CDI: 94 http://dx.doi.org/10.1128/mBio.01021-14 <strong>cdi_vincent_v3v5_results.tar.gz</strong> (<em>cdi_vincent</em>): H: 25, CDI: 25 http://dx.doi.org/10.1186/2049-2618-1-18 <strong>cdi_youngster_results.tar.gz</strong> (<em>cdi_youngster</em>): H: 4, CDI: 19 http://dx.doi.org/10.1093/cid/ciu135 <strong>crc_baxter_results.tar.gz</strong> (<em>crc_baxter</em>): adenoma: 198, H: 172, CRC: 120 http://dx.doi.org/10.1186/s13073-016-0290-3 <strong>crc_xiang_results.tar.gz</strong> (<em>crc_chen</em>): H: 22, CRC: 21 http://dx.doi.org/10.1371/journal.pone.0039743 <strong>crc_zackular_results.tar.gz</strong> (<em>crc_zackular</em>): adenoma: 30, H: 30, CRC: 30 http://dx.doi.org/10.1158/1940-6207.CAPR-14-0129 <strong>crc_zeller_results.tar.gz</strong> (<em>crc_zeller</em>): H: 75, CRC: 41 http://dx.doi.org/10.15252/msb.20145645 <strong>crc_zhao_results.tar.gz</strong> (<em>crc_wang</em>): H: 56, CRC: 46 http://dx.doi.org/10.1038/ismej.2011.109} <strong>edd_singh_results.tar.gz</strong> (<em>edd_singh</em>): STEC: 28, CAMP: 71, SALM: 66, SHIG: 34, H: 75 http://dx.doi.org/10.1186/s40168-015-0109-2 <strong>hiv_dinh_results.tar.gz</strong> (<em>hiv_dinh</em>): H: 16, HIV: 21 http://dx.doi.org/10.1093/infdis/jiu409 <strong>hiv_lozupone_results.tar.gz</strong> (<em>hiv_lozupone</em>): H: 13, HIV: 25 http://dx.doi.org/10.1016/j.chom.2013.08.006 <strong>hiv_noguerajulian_results.tar.gz</strong> (<em>hiv_noguerajulian</em>): H: 34, HIV: 206 https://doi.org/10.1016%2Fj.ebiom.2016.01.032 <strong>ibd_alm_results.tar.gz</strong> (<em>ibd_papa</em>): IBDundef: 1, nonIBD: 24, UC: 43, CD: 23 http://dx.doi.org/10.1371/journal.pone.0039242 <strong>ibd_engstrand_maxee_results.tar.gz</strong> (<em>ibd_willing</em>): CCD: 12, H: 35, ICD: 15, UC: 16, ICCD: 2 http://dx.doi.org/10.1053/j.gastro.2010.08.049 <strong>ibd_gevers_2014_results.tar.gz</strong> (<em>ibd_gevers</em>): H: 31, CD: 224 http://dx.doi.org/10.1016/j.chom.2014.02.005 <strong>ibd_huttenhower_results.tar.gz</strong> (<em>ibd_morgan</em>): H: 18, UC: 48, CD: 62 http://dx.doi.org/10.1186/gb-2012-13-9-r79 <strong>mhe_zhang_results.tar.gz</strong> (<em>liv_zhang</em>): CIRR: 25, H: 26, MHE: 26 http://dx.doi.org/10.1038/ajg.2013.221 <strong>nash_chan_results.tar.gz</strong> (<em>nash_wong</em>): H: 22, NASH: 16 http://dx.doi.org/10.1371/journal.pone.0062885 <strong>nash_ob_baker_results.tar.gz</strong> (<em>nash_zhu</em>): H: 16, NASH: 22, OB: 25 http://dx.doi.org/10.1002/hep.26093 <strong>ob_goodrich_results.tar.gz</strong> (<em>ob_goodrich</em>): OW: 322, H: 433, OB: 183 http://dx.doi.org/10.1016/j.cell.2014.09.053 <strong>ob_gordon_2008_v2_results.tar.gz</strong> (<em>ob_turnbaugh</em>): H: 61, OB: 219 http://dx.doi.org/10.1038/nature07540 <strong>ob_ross_results.tar.gz</strong> (<em>ob_ross</em>): H: 26, OB: 37 http://dx.doi.org/10.1186/s40168-015-0072-y <strong>ob_zupancic_results.tar.gz</strong> (<em>ob_zupancic</em>): H: 167, OB: 117 http://dx.doi.org/10.1371/journal.pone.0043052 <strong>par_scheperjans_results.tar.gz</strong> (<em>par_scheperjans</em>): H: 72, PAR: 72 http://dx.doi.org/10.1002/mds.26069 <strong>ra_littman_results.tar.gz</strong> (<em>art_scher</em>): H: 28, NORA: 44, CRA: 26, PSA: 16 http://dx.doi.org/10.7554/eLife.01202 <strong>t1d_alkanani_results.tar.gz</strong> (<em>t1d_alkanani</em>): T1D: 21, H: 55, T1D_new-onset: 35 http://dx.doi.org/10.2337/db14-1847 <strong>t1d_mejialeon_results.tar.gz</strong> (<em>t1d_mejialeon</em>): T1D: 21, H: 8 http://dx.doi.org/10.1038/srep03814 <strong>Version changes</strong> Changes in Version 2: added crc_zhu and ob_escobar datasets, as well as list of core genera.
### 概述 MicrobiomeHD是一个标准化的人类健康与疾病状态下肠道菌群研究数据库。该数据库纳入了已发表的病例-对照研究中的公开16S测序数据及其配套的患者元数据(metadata)。所有研究的原始测序数据均通过标准化流程完成下载与处理。 纳入MicrobiomeHD的数据集需满足以下条件: 1. 包含公开可用的原始测序数据(格式为fastq或fasta) 2. 包含公开可用的元数据,且每位患者至少标注有病例组与对照组分组信息 3. 病例组患者数量不少于15例 目前MicrobiomeHD的样本类型以粪便样本为主,部分数据集可能包含其他类型样本,具体信息可参见元数据。 ### 文件说明 本版MicrobiomeHD收录的数据集详细信息可参见GitHub仓库https://github.com/cduvallet/microbiomeHD中的`db/dataset_info.yaml`文件。顶级标识符对应Duvallet等人2017年研究中使用的数据集ID。YAML文件中列出的样本量为论文中报道的数值,可能与实际数据存在差异(例如存在缺失/额外数据、样本未通过质量控制等情况)。 所有数据集均通过标准化流程下载并处理,原始处理结果可通过本文档中的`*.tar.gz`文件获取。每个压缩包的目录结构与文件格式均遵循流程文档的规范:http://amplicon-sequencing-pipeline.readthedocs.io/en/latest/output.html。 重点关注的文件包括: - **"summary_file.txt"**:包含数据处理所用全部参数的汇总信息 - **"datasetID.metadata.txt"**:样本对应的元数据。需注意,元数据中部分样本可能未匹配到测序数据,反之亦然。 - **"RDP/datasetID.otu_table.100.denovo.rdp_assigned"**:使用RDP分类器(RDP classifier)赋予拉丁学名的100%相似度操作分类单元(Operational Taxonomic Unit, OTU)表。 - **"datasetID.otu_seqs.100.fasta"**:100%相似度OTU表中每个OTU的代表序列。OTU表中的OTU标签以`d__denovoID`结尾,该`denovoID`与本文件中的序列一一对应。 #### 处理流程说明 原始数据的获取方式详见Duvallet等人发表的论文《Meta analysis of microbiome studies identifies shared and disease-specific patterns》的补充材料。 原始测序数据通过Alm实验室自研的16S测序处理流程进行分析:https://github.com/thomasgurry/amplicon_sequencing_pipeline 该流程的官方文档可参见:http://amplicon-sequencing-pipeline.readthedocs.io/ 元数据均从原始论文或数据源中提取,并通过人工方式完成格式标准化。 ### 数据集贡献 MicrobiomeHD可用于提取单个病例-对照研究中的疾病特异性菌群信号。多数微生物对健康与疾病状态存在非特异性应答,单个研究中发现的绝大多数细菌关联均与该“核心”应答存在重叠。研究人员应将自身结果与本数据库的数据进行交叉验证,以确认所鉴定的微生物关联确实与所研究的疾病特异性相关。 本数据库提供了更新版的“核心”微生物列表,以及原始OTU表,供研究人员复现或调整分析方案以适配自身研究问题。 若您希望将自身的病例-对照数据集纳入MicrobiomeHD,请发送邮件至duvallet[at]mit.edu。 为确保我们能够通过标准流程处理您的数据,您需提供以下文件与信息: 1. fastq或fasta格式的原始测序数据(优先使用fastq格式) 2. 所需的处理步骤说明(例如去除引物或条形码、合并双端测序读段等) 3. 与测序数据对应的样本ID(可与序列中的条形码匹配,或对应每个已去多路复用的测序文件) 4. 每个样本的病例/对照分组元数据 5. 其他相关元数据(例如若样本非全部为粪便样本,请注明采样位点;若同一患者多次采样,请注明采样时间点等) 使用MicrobiomeHD进行分析即代表您同意将自身数据集贡献至本数据库,并承诺将原始测序数据(即fastq文件)公开共享。 ### 引用说明 MicrobiomeHD数据库及其收录的各数据集的原始文献详见Duvallet等人(2017)的研究:http://biorxiv.org/content/early/2017/05/08/134031 若您在分析中使用了本数据库中的任一数据集,请同时引用MicrobiomeHD(Duvallet等人(2017))以及对应数据集的原始发表文献。 Duvallet等人(2017)中用于数据处理与分析的代码可在GitHub仓库获取:https://github.com/cduvallet/microbiomeHD ### 补充文件 #### 核心微生物属 **"file-S3.core_genera.txt"**:Duvallet等人(2017)的补充表3,列出了与健康和疾病状态相关的核心微生物。 #### 数据集列表 需注意,MicrobiomeHD包含了Duvallet等人(2017)中的全部28个数据集,以及未满足该论文荟萃分析纳入标准的额外数据集。本版MicrobiomeHD收录的数据集详细信息可参见原始论文以及GitHub仓库https://github.com/cduvallet/microbiomeHD中的`db/dataset_info.yaml`文件。 本页面列出的样本量为原始论文中报道的数值,由于存在缺失数据、质量问题、条形码错配等情况,部分数值可能与实际数据存在差异。 - **"asd_son_results.tar.gz"**(<em>asd_son</em>):正常对照(NT):44例,孤独症谱系障碍(Autism Spectrum Disorder, ASD):59例 http://dx.doi.org/10.1371/journal.pone.0137725 - **"autism_kb_results.tar.gz"**(<em>asd_kang</em>):健康对照(H):20例,ASD:20例 http://dx.doi.org/10.1371/journal.pone.0068322 - **"cdi_schubert_results.tar.gz"**(<em>noncdi_schubert</em>):健康对照(H):155例,非艰难梭菌感染(nonCDI, Clostridioides difficile infection):89例,CDI:94例 http://dx.doi.org/10.1128/mBio.01021-14 - **"cdi_vincent_v3v5_results.tar.gz"**(<em>cdi_vincent</em>):健康对照(H):25例,CDI:25例 http://dx.doi.org/10.1186/2049-2618-1-18 - **"cdi_youngster_results.tar.gz"**(<em>cdi_youngster</em>):健康对照(H):4例,CDI:19例 http://dx.doi.org/10.1093/cid/ciu135 - **"crc_baxter_results.tar.gz"**(<em>crc_baxter</em>):腺瘤(adenoma):198例,健康对照(H):172例,结直肠癌(Colorectal Cancer, CRC):120例 http://dx.doi.org/10.1186/s13073-016-0290-3 - **"crc_xiang_results.tar.gz"**(<em>crc_chen</em>):健康对照(H):22例,CRC:21例 http://dx.doi.org/10.1371/journal.pone.0039743 - **"crc_zackular_results.tar.gz"**(<em>crc_zackular</em>):腺瘤:30例,健康对照(H):30例,CRC:30例 http://dx.doi.org/10.1158/1940-6207.CAPR-14-0129 - **"crc_zeller_results.tar.gz"**(<em>crc_zeller</em>):健康对照(H):75例,CRC:41例 http://dx.doi.org/10.15252/msb.20145645 - **"crc_zhao_results.tar.gz"**(<em>crc_wang</em>):健康对照(H):56例,CRC:46例 http://dx.doi.org/10.1038/ismej.2011.109 - **"edd_singh_results.tar.gz"**(<em>edd_singh</em>):产志贺毒素大肠杆菌(STEC):28例,弯曲菌感染(CAMP):71例,沙门菌感染(SALM):66例,志贺菌感染(SHIG):34例,健康对照(H):75例 http://dx.doi.org/10.1186/s40168-015-0109-2 - **"hiv_dinh_results.tar.gz"**(<em>hiv_dinh</em>):健康对照(H):16例,人类免疫缺陷病毒(HIV)感染:21例 http://dx.doi.org/10.1093/infdis/jiu409 - **"hiv_lozupone_results.tar.gz"**(<em>hiv_lozupone</em>):健康对照(H):13例,HIV感染:25例 http://dx.doi.org/10.1016/j.chom.2013.08.006 - **"hiv_noguerajulian_results.tar.gz"**(<em>hiv_noguerajulian</em>):健康对照(H):34例,HIV感染:206例 https://doi.org/10.1016%2Fj.ebiom.2016.01.032 - **"ibd_alm_results.tar.gz"**(<em>ibd_papa</em>):未明确分型的炎症性肠病(IBDundef, Inflammatory Bowel Disease):1例,非IBD对照:24例,溃疡性结肠炎(Ulcerative Colitis, UC):43例,克罗恩病(Crohn's Disease, CD):23例 http://dx.doi.org/10.1371/journal.pone.0039242 - **"ibd_engstrand_maxee_results.tar.gz"**(<em>ibd_willing</em>):慢性克罗恩病(CCD):12例,健康对照(H):35例,活动性克罗恩病(ICD):15例,UC:16例,慢性溃疡性结肠炎(ICCD):2例 http://dx.doi.org/10.1053/j.gastro.2010.08.049 - **"ibd_gevers_2014_results.tar.gz"**(<em>ibd_gevers</em>):健康对照(H):31例,CD:224例 http://dx.doi.org/10.1016/j.chom.2014.02.005 - **"ibd_huttenhower_results.tar.gz"**(<em>ibd_morgan</em>):健康对照(H):18例,UC:48例,CD:62例 http://dx.doi.org/10.1186/gb-2012-13-9-r79 - **"mhe_zhang_results.tar.gz"**(<em>liv_zhang</em>):肝硬化(CIRR):25例,健康对照(H):26例,轻微肝性脑病(Minimal Hepatic Encephalopathy, MHE):26例 http://dx.doi.org/10.1038/ajg.2013.221 - **"nash_chan_results.tar.gz"**(<em>nash_wong</em>):健康对照(H):22例,非酒精性脂肪性肝炎(Non-Alcoholic Steatohepatitis, NASH):16例 http://dx.doi.org/10.1371/journal.pone.0062885 - **"nash_ob_baker_results.tar.gz"**(<em>nash_zhu</em>):健康对照(H):16例,NASH:22例,肥胖(Obesity, OB):25例 http://dx.doi.org/10.1002/hep.26093 - **"ob_goodrich_results.tar.gz"**(<em>ob_goodrich</em>):超重(OW):322例,健康对照(H):433例,OB:183例 http://dx.doi.org/10.1016/j.cell.2014.09.053 - **"ob_gordon_2008_v2_results.tar.gz"**(<em>ob_turnbaugh</em>):健康对照(H):61例,OB:219例 http://dx.doi.org/10.1038/nature07540 - **"ob_ross_results.tar.gz"**(<em>ob_ross</em>):健康对照(H):26例,OB:37例 http://dx.doi.org/10.1186/s40168-015-0072-y - **"ob_zupancic_results.tar.gz"**(<em>ob_zupancic</em>):健康对照(H):167例,OB:117例 http://dx.doi.org/10.1371/journal.pone.0043052 - **"par_scheperjans_results.tar.gz"**(<em>par_scheperjans</em>):健康对照(H):72例,帕金森病(Parkinson's Disease, PAR):72例 http://dx.doi.org/10.1002/mds.26069 - **"ra_littman_results.tar.gz"**(<em>art_scher</em>):健康对照(H):28例,未分化类风湿关节炎(NORA):44例,活动性类风湿关节炎(CRA):26例,银屑病关节炎(PSA):16例 http://dx.doi.org/10.7554/eLife.01202 - **"t1d_alkanani_results.tar.gz"**(<em>t1d_alkanani</em>):1型糖尿病(Type 1 Diabetes, T1D):21例,健康对照(H):55例,新发T1D:35例 http://dx.doi.org/10.2337/db14-1847 - **"t1d_mejialeon_results.tar.gz"**(<em>t1d_mejialeon</em>):T1D:21例,健康对照(H):8例 http://dx.doi.org/10.1038/srep03814 ### 版本更新 V2版本更新内容:新增`crc_zhu`与`ob_escobar`数据集,以及核心微生物属列表。



