遇见数据集

Microbiomehd: The Human Gut Microbiome In Health And Disease

收藏
Zenodo2020-09-18 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

<strong>Overview</strong> MicrobiomeHD is a standardized database of human gut microbiome studies in health and disease. This database includes publicly available 16S data from published case-control studies and their associated patient metadata. Raw sequencing data for each study was downloaded and processed through a standardized pipeline. To be included in MicrobiomeHD, datasets have: publicly available raw sequencing data (fastq or fasta) publicly available metadata with at least case and control labels for each patient at least 15 case patients Currently, MicrobiomeHD is focused on stool samples. Additional samples may be included in certain datasets, as indicated in the metadata. <strong>Files</strong> Additional information about the datasets included in this MicrobiomeHD release are in the MicrobiomeHD github repo https://github.com/cduvallet/microbiomeHD, in the file <em>db/dataset_info.yaml</em>. Top-level identifiers correspond to the dataset IDs used in Duvallet et al. 2017. Sample sizes in the yaml file are those that were described in the papers, and may not exactly reflect the actual data (due to missing/extra data, samples which didn't pass quality control, etc). Each dataset was downloaded and processed through a standardized pipeline. The raw processing results are available in the *.tar.gz files here. Each file has the same directory structure and files, as described in the pipeline documentation: http://amplicon-sequencing-pipeline.readthedocs.io/en/latest/output.html. Specific files of interest include: <strong>summary_file.txt</strong>: this file contains a summary of all parameters used to process the data <strong>datasetID.metadata.txt</strong>: the metadata associated with the samples. Note that some samples in the metadata may not have sequencing data, and vice versa. <strong>RDP/datasetID.otu_table.100.denovo.rdp_assigned</strong>: the 100% OTU tables with Latin taxonomic names assigned using the RDP classifier (c = 0.5). <strong>datasetID.otu_seqs.100.fasta</strong>: representative sequences for each OTU in the 100% OTU table. OTU labels in the OTU table end with d__denovoID - these denovoIDs correspond to the sequences in this file. The raw data was acquired as described in the supplementary materials of Duvallet et al.'s "Meta analysis of microbiome studies identifies shared and disease-specific patterns". Raw sequencing data was processed with the Alm lab's in-house 16S processing pipeline: https://github.com/thomasgurry/amplicon_sequencing_pipeline Pipeline documentation is available at: http://amplicon-sequencing-pipeline.readthedocs.io/ Metadata was extracted from the original papers and/or data sources, and formatted manually. <strong>Contributing</strong> MicrobiomeHD is a resource that can be used to extract disease-specific microbiome signals in individual case-control studies. Many microbes respond non-specifically to health and disease, and the majority of bacterial associations within individual studies overlap with this "core" response. Researchers should cross-check their results with the data presented here to ensure that their identified microbial associations are specific to their disease under study. We provide an updated list of "core" microbes here, as well as the raw OTU tables for anyone who wishes to reproduce and adapt this analysis to their study question. If you would like to include your case-control dataset in MicrobiomeHD, please email duvallet[at]mit.edu. For us to process your data through our standard pipeline, you will need to provide the following files and information about your data: raw sequencing data in fastq or fasta format (preferably fastq) information about which processing steps will be required (e.g. removing primers or barcodes, merging paired-end reads, etc) sample IDs associated with the sequencing data (either mapped to barcodes still in the sequences, or to each de-multiplexed sequencing file) case/control metadata of each sample other relevant metadata (e.g. sampling site, if not all samples are stool; sampling time point, if multiple samples per patient were taken; etc) By using MicrobiomeHD in your own analyses, you agree to contribute your dataset to this database and to make your raw sequencing data (i.e. fastq files) publicly available. <strong>Citing MicrobiomeHD</strong> The MicrobiomeHD database and original publications for each of these datasets are described in Duvallet et al. (2017): http://biorxiv.org/content/early/2017/05/08/134031 If you use any of these datasets in your analysis, please cite both MicrobiomeHD (Duvallet et al. (2017)) and the original publication for each dataset that you use. The code used to process and analyze this data in Duvallet et al. (2017) is available on github: https://github.com/cduvallet/microbiomeHD <strong>Files</strong> <em>Data files</em> <strong>file-S3.core_genera.txt</strong>: Supplemental Table 3 from Duvallet et al. (2017), listing the core health- and disease-associated microbes.<br> <strong>dataset_info.yaml</strong>: yaml file with additional dataset metadata. <em>Datasets</em> Note that MicrobiomeHD contains all 28 datasets from Duvallet et al. (2017), as well as additional datasets which did not meet the inclusion criteria for the meta-analysis presented in the paper. Additional information about the datasets included in this MicrobiomeHD release are in the original publications and the MicrobiomeHD github repo https://github.com/cduvallet/microbiomeHD, and in the file <em>dataset_info.yaml</em>. The sample sizes listed here reflect what was reported in the original publications. Some may have discrepancies between what is reported and what is in the actual data due to missing data, quality issues, barcode mismatches, etc. <strong>asd_son_results.tar.gz</strong> (<em>asd_son</em>): NT: 44, ASD: 59 http://dx.doi.org/10.1371/journal.pone.0137725 <strong>autism_kb_results.tar.gz</strong> (<em>asd_kang</em>): H: 20, ASD: 20 http://dx.doi.org/10.1371/journal.pone.0068322 <strong>cdi_schubert_results.tar.gz</strong> (<em>noncdi_schubert</em>): H: 155, nonCDI: 89, CDI: 94 http://dx.doi.org/10.1128/mBio.01021-14 <strong>cdi_vincent_v3v5_results.tar.gz</strong> (<em>cdi_vincent</em>): H: 25, CDI: 25 http://dx.doi.org/10.1186/2049-2618-1-18 <strong>cdi_youngster_results.tar.gz</strong> (<em>cdi_youngster</em>): H: 4, CDI: 19 http://dx.doi.org/10.1093/cid/ciu135 <strong>crc_baxter_results.tar.gz</strong> (<em>crc_baxter</em>): adenoma: 198, H: 172, CRC: 120 http://dx.doi.org/10.1186/s13073-016-0290-3 <strong>crc_xiang_results.tar.gz</strong> (<em>crc_chen</em>): H: 22, CRC: 21 http://dx.doi.org/10.1371/journal.pone.0039743 <strong>crc_zackular_results.tar.gz</strong> (<em>crc_zackular</em>): adenoma: 30, H: 30, CRC: 30 http://dx.doi.org/10.1158/1940-6207.CAPR-14-0129 <strong>crc_zeller_results.tar.gz</strong> (<em>crc_zeller</em>): H: 75, CRC: 41 http://dx.doi.org/10.15252/msb.20145645 <strong>crc_zhao_results.tar.gz</strong> (<em>crc_wang</em>): H: 56, CRC: 46 http://dx.doi.org/10.1038/ismej.2011.109} <strong>edd_singh_results.tar.gz</strong> (<em>edd_singh</em>): STEC: 28, CAMP: 71, SALM: 66, SHIG: 34, H: 75 http://dx.doi.org/10.1186/s40168-015-0109-2 <strong>hiv_dinh_results.tar.gz</strong> (<em>hiv_dinh</em>): H: 16, HIV: 21 http://dx.doi.org/10.1093/infdis/jiu409 <strong>hiv_lozupone_results.tar.gz</strong> (<em>hiv_lozupone</em>): H: 13, HIV: 25 http://dx.doi.org/10.1016/j.chom.2013.08.006 <strong>hiv_noguerajulian_results.tar.gz</strong> (<em>hiv_noguerajulian</em>): H: 34, HIV: 206 https://doi.org/10.1016%2Fj.ebiom.2016.01.032 <strong>ibd_alm_results.tar.gz</strong> (<em>ibd_papa</em>): IBDundef: 1, nonIBD: 24, UC: 43, CD: 23 http://dx.doi.org/10.1371/journal.pone.0039242 <strong>ibd_engstrand_maxee_results.tar.gz</strong> (<em>ibd_willing</em>): CCD: 12, H: 35, ICD: 15, UC: 16, ICCD: 2 http://dx.doi.org/10.1053/j.gastro.2010.08.049 <strong>ibd_gevers_2014_results.tar.gz</strong> (<em>ibd_gevers</em>): H: 31, CD: 224 http://dx.doi.org/10.1016/j.chom.2014.02.005 <strong>ibd_huttenhower_results.tar.gz</strong> (<em>ibd_morgan</em>): H: 18, UC: 48, CD: 62 http://dx.doi.org/10.1186/gb-2012-13-9-r79 <strong>mhe_zhang_results.tar.gz</strong> (<em>liv_zhang</em>): CIRR: 25, H: 26, MHE: 26 http://dx.doi.org/10.1038/ajg.2013.221 <strong>nash_chan_results.tar.gz</strong> (<em>nash_wong</em>): H: 22, NASH: 16 http://dx.doi.org/10.1371/journal.pone.0062885 <strong>nash_ob_baker_results.tar.gz</strong> (<em>nash_zhu</em>): H: 16, NASH: 22, OB: 25 http://dx.doi.org/10.1002/hep.26093 <strong>ob_goodrich_results.tar.gz</strong> (<em>ob_goodrich</em>): OW: 322, H: 433, OB: 183 http://dx.doi.org/10.1016/j.cell.2014.09.053 <strong>ob_gordon_2008_v2_results.tar.gz</strong> (<em>ob_turnbaugh</em>): H: 61, OB: 219 http://dx.doi.org/10.1038/nature07540 <strong>ob_ross_results.tar.gz</strong> (<em>ob_ross</em>): H: 26, OB: 37 http://dx.doi.org/10.1186/s40168-015-0072-y <strong>ob_zupancic_results.tar.gz</strong> (<em>ob_zupancic</em>): H: 167, OB: 117 http://dx.doi.org/10.1371/journal.pone.0043052 <strong>par_scheperjans_results.tar.gz</strong> (<em>par_scheperjans</em>): H: 72, PAR: 72 http://dx.doi.org/10.1002/mds.26069 <strong>ra_littman_results.tar.gz</strong> (<em>art_scher</em>): H: 28, NORA: 44, CRA: 26, PSA: 16 http://dx.doi.org/10.7554/eLife.01202 <strong>t1d_alkanani_results.tar.gz</strong> (<em>t1d_alkanani</em>): T1D: 21, H: 55, T1D_new-onset: 35 http://dx.doi.org/10.2337/db14-1847 <strong>t1d_mejialeon_results.tar.gz</strong> (<em>t1d_mejialeon</em>): T1D: 21, H: 8 http://dx.doi.org/10.1038/srep03814 <strong>Version changes</strong> Changes in Version 2: added crc_zhu and ob_escobar datasets, as well as list of core genera and dataset_info.yaml.

<strong>概述</strong> MicrobiomeHD是一个面向健康与疾病状态下人类肠道微生物组研究的标准化数据库。该库收录了已发表病例对照研究中的公开16S测序数据,以及对应的患者元数据(metadata)。所有研究的原始测序数据均经标准化流程下载并处理。 纳入MicrobiomeHD的数据集需满足以下要求: - 具备公开可获取的原始测序数据(格式为fastq或fasta); - 附带公开元数据,且每个样本至少包含病例组与对照组的分组标签; - 病例组患者数量不少于15例。 目前MicrobiomeHD的数据集以粪便样本为主,部分数据集可能包含其他类型样本,具体信息可参见元数据。 <strong>文件</strong> 本版本MicrobiomeHD所包含数据集的详细信息,可访问其GitHub仓库https://github.com/cduvallet/microbiomeHD中的<em>db/dataset_info.yaml</em>文件获取。数据集的顶层标识符对应Duvallet等人2017年研究中使用的数据集ID。YAML文件中列出的样本量为原论文中报道的数值,可能与实际数据存在偏差(例如存在数据缺失、额外数据或样本未通过质量质控等情况)。 所有数据集均经标准化流程下载并处理,原始处理结果以*.tar.gz格式文件提供。每个压缩包的目录结构与文件组成均保持一致,具体可参考流程文档:http://amplicon-sequencing-pipeline.readthedocs.io/en/latest/output.html。 需关注的特定文件包括: <strong>summary_file.txt</strong>:该文件包含数据处理所用全部参数的汇总信息。 <strong>datasetID.metadata.txt</strong>:对应样本的元数据。需注意,元数据中列出的部分样本可能并无对应的测序数据,反之亦然。 <strong>RDP/datasetID.otu_table.100.denovo.rdp_assigned</strong>:采用RDP分类器(RDP classifier,置信度阈值c=0.5)赋予拉丁分类学名称的100%聚类操作分类单元(Operational Taxonomic Unit, OTU)表。 <strong>datasetID.otu_seqs.100.fasta</strong>:100%聚类OTU表中每个OTU的代表序列。OTU表中的OTU标签以d__denovoID结尾,这些denovoID与本文件中的序列一一对应。 原始数据的获取方式详见Duvallet等人发表的论文《Meta analysis of microbiome studies identifies shared and disease-specific patterns》的补充材料。 原始测序数据由Alm实验室自研的16S测序处理流程完成分析:https://github.com/thomasgurry/amplicon_sequencing_pipeline 该流程的文档可访问:http://amplicon-sequencing-pipeline.readthedocs.io/ 元数据均从原始论文或数据源中提取,并经手动格式化整理。 <strong>数据共建</strong> MicrobiomeHD可用于从单病例对照研究中提取疾病特异性微生物组信号。多数微生物对健康与疾病状态的响应不具有疾病特异性,单个研究中发现的大部分细菌关联均属于这种“核心”响应范畴。研究者应将自身研究结果与本数据库中的数据进行交叉验证,以确认所识别的微生物关联确实对应其所研究的特定疾病。 本库提供了更新后的“核心”微生物列表,同时也提供了原始OTU表,以供有需要的研究者复现该分析或将其适配至自身研究课题中。 若您希望将自己的病例对照数据集纳入MicrobiomeHD,请发送邮件至duvallet[at]mit.edu。 若需通过本库的标准化流程处理您的数据,请提供以下文件及数据相关信息: - 格式为fastq或fasta的原始测序数据(优先提供fastq格式); - 所需的处理步骤说明(例如移除引物或接头、合并双端测序读段等); - 与测序数据对应的样本ID(可与序列中的条形码绑定,或对应每个已去多路复用的测序文件); - 每个样本的病例/对照组分组元数据; - 其他相关元数据(例如若样本非全部为粪便样本,请说明采样位点;若同一患者有多个采样时间点,请说明采样时间等)。 若您在自身分析中使用MicrobiomeHD,则视为同意将您的数据集贡献至本数据库,并将原始测序数据(即fastq文件)公开获取。 <strong>引用说明</strong> MicrobiomeHD数据库及各数据集的原始发表文献详见Duvallet等人2017年的研究:http://biorxiv.org/content/early/2017/05/08/134031 若您在分析中使用本库中的任一数据集,请同时引用MicrobiomeHD(Duvallet等人2017年)以及您所使用的各数据集对应的原始发表文献。 Duvallet等人2017年研究中用于数据处理与分析的代码已上传至GitHub:https://github.com/cduvallet/microbiomeHD <strong>文件</strong> <em>数据文件</em> <strong>file-S3.core_genera.txt</strong>:Duvallet等人2017年研究的补充表3,列出了与健康及疾病状态相关的核心微生物属。<br> <strong>dataset_info.yaml</strong>:包含额外数据集元数据的YAML格式文件。 <em>数据集列表</em> 需注意,MicrobiomeHD不仅包含Duvallet等人2017年研究中的全部28个数据集,还收录了未达到该论文荟萃分析纳入标准的额外数据集。本版本MicrobiomeHD所包含数据集的详细信息,可参见原始发表文献、本库的GitHub仓库https://github.com/cduvallet/microbiomeHD,以及<em>dataset_info.yaml</em>文件。 此处列出的样本量为原论文中报道的数值,可能与实际数据存在偏差(例如存在数据缺失、质量问题、条形码错配等情况)。 <strong>asd_son_results.tar.gz</strong>(<em>asd_son</em>):典型发育对照(NT)44例,孤独症谱系障碍(ASD)59例 http://dx.doi.org/10.1371/journal.pone.0137725 <strong>autism_kb_results.tar.gz</strong>(<em>asd_kang</em>):健康对照(H)20例,孤独症谱系障碍(ASD)20例 http://dx.doi.org/10.1371/journal.pone.0068322 <strong>cdi_schubert_results.tar.gz</strong>(<em>noncdi_schubert</em>):健康对照(H)155例,非艰难梭菌感染(nonCDI)89例,艰难梭菌感染(CDI)94例 http://dx.doi.org/10.1128/mBio.01021-14 <strong>cdi_vincent_v3v5_results.tar.gz</strong>(<em>cdi_vincent</em>):健康对照(H)25例,艰难梭菌感染(CDI)25例 http://dx.doi.org/10.1186/2049-2618-1-18 <strong>cdi_youngster_results.tar.gz</strong>(<em>cdi_youngster</em>):健康对照(H)4例,艰难梭菌感染(CDI)19例 http://dx.doi.org/10.1093/cid/ciu135 <strong>crc_baxter_results.tar.gz</strong>(<em>crc_baxter</em>):腺瘤198例,健康对照(H)172例,结直肠癌(CRC)120例 http://dx.doi.org/10.1186/s13073-016-0290-3 <strong>crc_xiang_results.tar.gz</strong>(<em>crc_chen</em>):健康对照(H)22例,结直肠癌(CRC)21例 http://dx.doi.org/10.1371/journal.pone.0039743 <strong>crc_zackular_results.tar.gz</strong>(<em>crc_zackular</em>):腺瘤30例,健康对照(H)30例,结直肠癌(CRC)30例 http://dx.doi.org/10.1158/1940-6207.CAPR-14-0129 <strong>crc_zeller_results.tar.gz</strong>(<em>crc_zeller</em>):健康对照(H)75例,结直肠癌(CRC)41例 http://dx.doi.org/10.15252/msb.20145645 <strong>crc_zhao_results.tar.gz</strong>(<em>crc_wang</em>):健康对照(H)56例,结直肠癌(CRC)46例 http://dx.doi.org/10.1038/ismej.2011.109 <strong>edd_singh_results.tar.gz</strong>(<em>edd_singh</em>):肠出血性大肠杆菌(STEC)28例,弯曲杆菌(CAMP)71例,沙门氏菌(SALM)66例,志贺氏菌(SHIG)34例,健康对照(H)75例 http://dx.doi.org/10.1186/s40168-015-0109-2 <strong>hiv_dinh_results.tar.gz</strong>(<em>hiv_dinh</em>):健康对照(H)16例,人类免疫缺陷病毒感染(HIV)21例 http://dx.doi.org/10.1093/infdis/jiu409 <strong>hiv_lozupone_results.tar.gz</strong>(<em>hiv_lozupone</em>):健康对照(H)13例,人类免疫缺陷病毒感染(HIV)25例 http://dx.doi.org/10.1016/j.chom.2013.08.006 <strong>hiv_noguerajulian_results.tar.gz</strong>(<em>hiv_noguerajulian</em>):健康对照(H)34例,人类免疫缺陷病毒感染(HIV)206例 https://doi.org/10.1016%2Fj.ebiom.2016.01.032 <strong>ibd_alm_results.tar.gz</strong>(<em>ibd_papa</em>):未明确分型的炎症性肠病(IBDundef)1例,非IBD对照24例,溃疡性结肠炎(UC)43例,克罗恩病(CD)23例 http://dx.doi.org/10.1371/journal.pone.0039242 <strong>ibd_engstrand_maxee_results.tar.gz</strong>(<em>ibd_willing</em>):慢性克罗恩病(CCD)12例,健康对照(H)35例,慢性溃疡性结肠炎(ICD)15例,溃疡性结肠炎(UC)16例,初发性克罗恩病(ICCD)2例 http://dx.doi.org/10.1053/j.gastro.2010.08.049 <strong>ibd_gevers_2014_results.tar.gz</strong>(<em>ibd_gevers</em>):健康对照(H)31例,克罗恩病(CD)224例 http://dx.doi.org/10.1016/j.chom.2014.02.005 <strong>ibd_huttenhower_results.tar.gz</strong>(<em>ibd_morgan</em>):健康对照(H)18例,溃疡性结肠炎(UC)48例,克罗恩病(CD)62例 http://dx.doi.org/10.1186/gb-2012-13-9-r79 <strong>mhe_zhang_results.tar.gz</strong>(<em>liv_zhang</em>):肝硬化(CIRR)25例,健康对照(H)26例,轻微肝性脑病(MHE)26例 http://dx.doi.org/10.1038/ajg.2013.221 <strong>nash_chan_results.tar.gz</strong>(<em>nash_wong</em>):健康对照(H)22例,非酒精性脂肪性肝炎(NASH)16例 http://dx.doi.org/10.1371/journal.pone.0062885 <strong>nash_ob_baker_results.tar.gz</strong>(<em>nash_zhu</em>):健康对照(H)16例,非酒精性脂肪性肝炎(NASH)22例,肥胖(OB)25例 http://dx.doi.org/10.1002/hep.26093 <strong>ob_goodrich_results.tar.gz</strong>(<em>ob_goodrich</em>):超重(OW)322例,健康对照(H)433例,肥胖(OB)183例 http://dx.doi.org/10.1016/j.cell.2014.09.053 <strong>ob_gordon_2008_v2_results.tar.gz</strong>(<em>ob_turnbaugh</em>):健康对照(H)61例,肥胖(OB)219例 http://dx.doi.org/10.1038/nature07540 <strong>ob_ross_results.tar.gz</strong>(<em>ob_ross</em>):健康对照(H)26例,肥胖(OB)37例 http://dx.doi.org/10.1186/s40168-015-0072-y <strong>ob_zupancic_results.tar.gz</strong>(<em>ob_zupancic</em>):健康对照(H)167例,肥胖(OB)117例 http://dx.doi.org/10.1371/journal.pone.0043052 <strong>par_scheperjans_results.tar.gz</strong>(<em>par_scheperjans</em>):健康对照(H)72例,帕金森病(PAR)72例 http://dx.doi.org/10.1002/mds.26069 <strong>ra_littman_results.tar.gz</strong>(<em>art_scher</em>):健康对照(H)28例,未分化结缔组织病(NORA)44例,类风湿关节炎(CRA)26例,银屑病关节炎(PSA)16例 http://dx.doi.org/10.7554/eLife.01202 <strong>t1d_alkanani_results.tar.gz</strong>(<em>t1d_alkanani</em>):1型糖尿病(T1D)21例,健康对照(H)55例,新发1型糖尿病(T1D_new-onset)35例 http://dx.doi.org/10.2337/db14-1847 <strong>t1d_mejialeon_results.tar.gz</strong>(<em>t1d_mejialeon</em>):1型糖尿病(T1D)21例,健康对照(H)8例 http://dx.doi.org/10.1038/srep03814 <strong>版本更新</strong> 版本2更新内容:新增crc_zhu与ob_escobar数据集,同时补充了核心微生物属列表及dataset_info.yaml文件。

提供机构:
Zenodo
创建时间:
2017-08-08
二维码
社区交流群
二维码
科研交流群
商业服务