遇见数据集

Supplementary Data for "Sugar-rich foods exacerbate antibiotic-induced microbiome injury"

收藏
Zenodo2026-07-31 更新2026-08-02 收录
官方服务:

资源简介:

Data S1. Sequencing Data Accession Numbers and Metadata for Microbiome and Host Gene Expression Samples This file contains five sheets detailing the sample identifiers, experimental metadata, and NCBI Sequence Read Archive (SRA) accession numbers for the study. Every sheet lists the BioProject each run belongs to, in addition to the run accession. Sheet 1: Human 16S rRNA Amplicon Sequencing. A mapping of sample IDs to their corresponding SRA run accession numbers for human microbiome amplicon data, with the BioProject each run belongs to. The 1,009 samples are distributed across six BioProjects: PRJNA878528 (476 samples), PRJNA545312 (484), PRJNA810900 (38), PRJNA972747 (5), PRJNA979931 (5) and PRJNA810891 (1). Sheet 2: Human Shotgun Metagenomics. Detailed mapping for human shotgun sequencing samples, including sample IDs, patient IDs (PID), and corresponding SRA, BioProject, and BioSample accession numbers. Sheet 3: Mouse 16S rRNA Amplicon Sequencing. Metadata for the 111 mouse fecal amplicon samples (SRR31624460–SRR31624570, BioProject PRJNA878528): sample IDs, experiment number, collection timepoint (day 1, 3 or 6), mouse identifier, and the antibiotic and diet treatments. Sheet 4: Mouse Shallow Shotgun Metagenomics. Metadata for the 24 mouse shallow shotgun samples, listing SRA run accessions (SRR37009193–SRR37009216, BioProject PRJNA878528), BioSample accessions, library IDs, timepoints, treatments, sample source and the sequencing library details. Sheet 5: Mouse RNA-seq. Metadata for host transcriptomic profiling of mouse intestinal epithelium, 20 libraries in total: 10 large intestine (SRR36928799–SRR36928808) and 10 small intestine (SRR39866831–SRR39866840), BioProject PRJNA878528. Includes mouse IDs, gavage and diet treatments, BioSample accessions and the sequencing library details. Data S2. Detailed Nutritional Intake Data for Anonymized Patients (filename: 152_combined_DTB.csv) This file contains comprehensive, anonymized data on patient dietary intake. Columns include: pid: Patient ID. Meal: Meal category (e.g., breakfast, lunch, dinner). Food_NSC: Food name. fdrt: Diet entry day, relative to transplant. Unit: Unit of measurement for food quantity (e.g., grams, ounces). Por_eaten: Portion of food consumed. Food_code, description: Food code and corresponding description from the Food and Nutrient Database for Dietary Studies (FNDDS). Total calories, weight: Caloric content and weight of the consumed portion. Individual macronutrients (grams): Gram weight of each macronutrient (e.g., protein, fat, sugar, fiber as well as carbohydrate that excludes sugar and fiber) in the consumed portion. Dehydrated weight: Total weight of the consumed portion minus the water weight. Data S3. Summarized Dietary Intake, Clinical Variables, and Macronutrients for Bayesian Modeling (filename: 153_combined_META.csv) This file provides summarized dietary intake data, macronutrient profiles, and relevant clinical variables used in the Bayesian model. Columns include: pid: Unique patient identifier. sampleid: Unique stool sample identifier. sdrt: Stool sample collection day, relative to transplant. fg_egg ... fg_veggie: Average intake (in grams) of foods belonging to nine broad food groups (e.g., eggs, vegetables) during the two days preceding stool sample collection. ave_Calories_kcal, ave_Protein_g, ave_Fat_g, ave_Carbohydrates_g, ave_Fibers_g, ave_Sugars_g: Average daily caloric intake (kcal) and macronutrient intake (grams of protein, fat, carbohydrates, fiber, and sugar) during the two days prior to stool sample collection. intensity: Intensity of the conditioning regimen. empirical: Binary indicator (TRUE/FALSE) of patient exposure to specific antibiotics (piperacillin/tazobactam, carbapenems, cefepime, linezolid, oral vancomycin, and metronidazole) in the two days prior to stool sample collection. simpson_reciprocal: α-diversity of the stool sample, calculated using the Simpson reciprocal index. TPN, EN: Binary indicators (TRUE/FALSE) of a patient receiving total parenteral nutrition (TPN) or enteral nutrition (EN) in the two days prior to stool sample collection. timebin: Time interval of stool sample collection, categorized by week relative to transplant. Data S4. Medication Exposure Overlapping With Stool Samples This file contains a record of all medication exposures that occurred during the 48-hour period before each stool sample was collected. This window was chosen to investigate the potential impact of recent medication use on the stool microbiome. The table includes the following columns: sampleid: A unique identifier assigned to each stool sample. pid: Patient ID. sdrt: Stool sample collection day, relative to transplant. class: The pharmacological class of the administered medication (e.g., "quinolones", "beta-lactamase inhibitors", etc.). drug_name_clean: The name of the medication (e.g., "ciprofloxacin", "vancomycin"). route_clean: The route of administration (e.g., "IV", "oral"). drug_category_for_this_study: A study-specific categorization of the medication, based on its potential impact on the gut microbiome. The categories are: > * broad_spectrum: Broad-spectrum antibiotics, as classified in this study: piperacillin/tazobactam, carbapenems, cefepime, linezolid, oral vancomycin, and metronidazole > * fluoroquinolones: Fluoroquinolone antibiotics (ciprofloxacin or levofloxacin). > * other_antibacterial: Antibacterial medications not classified as broad-spectrum or fluoroquinolones. > * not_antibacterial: Medications not expected to have a direct antibacterial effect. Data S5. Nutrient Profiles of Foods Consumed (FNDDS 2015-2016) This supplemental file contains the complete nutrient values for food codes listed in the USDA Food and Nutrient Database for Dietary Studies (FNDDS) 2015-2016. All nutrient values in this file are reported per 100 grams of edible portion, meaning it includes the water content of the food. See the column labeled "Water (g)" explicitly lists the grams of water contained within each 100g portion. This table is downloaded from: https://www.ars.usda.gov/northeast-area/beltsville-md-bhnrc/beltsville-human-nutrition-research-center/food-surveys-research-group/docs/fndds-download-databases/ Data S6. Individual patient data. Line plots illustrate the observed patterns in daily caloric intake, daily diet α-diversity (Faith's phylogenetic distance), and fecal microbiota α-diversity (inverse Simpson index) across patient hospitalizations. Each panel displays the longitudinal data for a single patient. The number of days with recorded dietary data and the number of evaluable stool samples are annotated in red and blue, respectively, within each panel. The red, black, and blue lines plot daily caloric intake, diet α-diversity, and fecal microbiota α-diversity, respectively. Patients with only one or no stool sample collected will not have a line representing their fecal α-diversity over time. The first two have values on the same numeric scale and share the same left y-axis. Microbiome α-diversity value is quantified on the y-axis on the right. The vertical grey line is the day of cell infusion (day 0). The vertical dotted green line is the day of neutrophil engraftment. All panels share the same y-axis ranges; x-axis ranges are specific to the per-patient data availability. Alphanumeric codes are anonymized patient identifiers. Data S7. Source data for the mouse experiments — Dai_mouse_figure_raw_data.xlsx Sixteen worksheets, one per published mouse panel plus the two sequencing tables. This is the source-data workbook for Figure 4c and Extended Data Figures 8 and 9. Mice are identified only by internal cage and mouse numbers. Colony-forming unit (CFU) experiments. Unless stated otherwise, values are Enterococcus CFU per milligram of stool, counted on selective agar; multiply by 1000 for CFU per gram. Each fecal pellet was weighed, resuspended in 1 mL of buffer, and a dilution series plated at a fixed volume. Figure_4c_indiv_days_data — CFU per milligram for each mouse on each day (Treatment + Day, CFUs per mg stool, Treatment Group, Experiment). Figure_4c_trapezoidal_auc — per-mouse area under the CFU-versus-day curve, integrated on the linear scale over days 1, 3 and 6 (Treatment Group, Trapezoidal AUC (CFU per mg stool x days), Experiment). Sup_Figure_8j_delayed_sucrose_a — area under the log10 CFU curve for the delayed-sucrose experiment, with the experiment number so replicates can be separated (Treatment Group, AUC of log10 (CFU per mg stool) x days, Experiment Number). Sup_Figure_8k_delayed_sucrose_C — the same experiment as raw per-mouse longitudinal counts (experiment_no, mouse_identifier, day, treatment, cfus_per_mg_stool). Sup_Figure_8l_no_fiber_chow_CFU — mice on fiber-free chow; values are already log10-transformed (Treatment + Day, log10 (Enterococcal CFU per mg stool), Treatment Group). Sup_Figure_9a_21_day_exp — the 21-day vehicle-versus-sucrose experiment, keyed by a single combined abx_treatment__diet_treatment__day label (abx_treatment__diet_treatment__day, cfus_per_mg_stool). Sup_Figure_9c_alternate_sugars — the four-sugar comparison under antibiotic exposure, same two columns and key format. Sup_Figure_9b_smoothie — the smoothie diet experiment, with per-mouse identifiers (Treatement + Day, CFUs per mg stool, Treatment, Treatment + Mouse Number, Experiment; the first header is misspelled in the file and is left as it is so the analysis scripts keep matching it). Sup_Figure_9d_time_course_monoc — germ-free mono-colonization time course, in hours post-inoculation rather than days (Time post innoculation (hours), CFUs per mg stool, Mouse Identifier, Treatment). Sup_Figure_9h_bmt_cfu — bone-marrow-transplant recipients; includes both the T-cell-depleted (bone marrow only) and bone marrow plus T-cell arms (Treatment + Day, CFUs per mg stool, Treatment Group). Chow consumption and body weight. Sup_Figure_8g_chow_consumed_day — grams of chow consumed per day (treatment, chow_per_day, day). Sup_Figure_8h_chow_consumed_auc — the same, as area under the curve. Sup_Figure_8m_all_mouse_weight_ — per-mouse body weight and percentage change from baseline (treatment_group, day, percent_change, mice_weights, experiment_no, mouse_identifier). Sequencing tables. mouse_16s_asv_relab — mouse fecal 16S relative abundances: rows are amplicon sequence variants labelled seqNNNN:lineage in the Taxon column, the remaining 111 columns are samples (PL0109 and so on), and values are relative abundances. mouse_16s_sample_meta — one row per 16S sample: sampleid, experiment_no, day (1, 3 or 6), group, mouse_no, tube_no, investigator, abx_treatment (PBS or biapenem), diet_treatment (vehicle or sucrose) and simpson_reciprocal, the inverse Simpson alpha diversity. raw_counts_matrix — gene-level read counts for the 20 RNA-seq libraries, generated with featureCounts against the GENCODE mouse basic gene annotation release vM37 (GRCm39, equivalent to Ensembl 114). Column 1 is Geneid, a versioned Ensembl mouse identifier; the other 20 columns are libraries named {experiment}_{mouse}_{segment}_IGO_{project}_{library}, where the segment is LI (large intestine) or SI (small intestine). Additional released tables. These complete the set needed to regenerate the publicly reproducible figures. They are identical to the released_data/ folder of the analysis code repository, so either source may be used. All are de-identified, and patient identifiers are the same arbitrary pid codes used in Data S2 and S3. Two time conventions run through the whole deposit: fdrt is a diet record's day relative to transplant and sdrt a stool sample's day relative to transplant, with transplant as day 0 and negative values before it. Processed human fecal 16S tables. 63_asv_count_relab_res.csv (104,969 rows) — per-sample amplicon sequence variant abundances. asv_key is the variant identifier, sampleid the stool sample, count the read count and count_relative the within-sample relative abundance. 63_asv_blast_annotation.csv (14,430 rows, one per variant) — taxonomic assignment for each variant and the BLAST hit that produced it. kingdom through species give the lineage; query_length is the length of the variant sequence, align_length the length of the alignment, pident the percent identity, nident the number of identical bases, score the raw alignment score, and length_ratio the alignment length divided by the query length, so a value near 1 indicates a full-length match. 45_quality_asv_relab_pident97_genus.csv (104,969 rows) — the same variant abundances with a genus attached, but only where the BLAST identity supports it: genus is set to missing wherever pident is at or below 97, so genus-level summaries can be restricted to confident assignments. Processed human fecal shotgun table. mgx_enterococcus_species_relab.csv (429 rows) — Enterococcus relative abundance resolved to species from the shotgun metagenomes: sample, species and relab, in long format with one row per sample-species pair. Derived analysis inputs. R59_meta_expanded.csv (1,009 rows) — Data S3 with five columns appended for the taxon abundance models: ci_cleaned_numeric (hematopoietic cell transplantation comorbidity index), disease_lineage (myeloid or lymphoid), PCA (patient-controlled analgesia, used as a proxy for mucositis severity; true if the patient was receiving it in the two days before stool collection, false otherwise), exposure_type (the reason for broad-spectrum antibiotic exposure in that same two-day window: empiric, bsi_treatment for pathogen-directed treatment of bloodstream infection, cdiff_treatment for Clostridioides difficile diarrhea, or no_broad_spectrum_exposure) and asv_1_clr, the centered-log-ratio abundance of amplicon sequence variant 1, Enterococcus faecium. taxumap_embedding.csv (4,081 rows) — the taxonomy-weighted UMAP embedding of daily diet records. index_column is a patient-day key formed as P followed by the patient identifier, an underscore, and fdrt (so PP1_-6 is patient P1, six days before transplant); taxumap1 and taxumap2 are the two embedding coordinates. diet-alpha-diversity.tsv (4,081 rows) — daily dietary alpha diversity as Faith's phylogenetic diversity computed on the food taxonomy tree. The first column is an unnamed patient-day key formed as the patient identifier, the letter d, then fdrt (so P100d-1 is patient P100 on the day before transplant); faith_pd is the value. 072_total_patients_zero_eating_days_pid.csv (98 rows) — patient-days on which a participant was confirmed to have eaten nothing (pid, fdrt, diet_data_status). Food taxonomy tree files (used to build Figure 1d and the diet diversity metrics). NodeLabelsMCT.txt (1,479 rows) — the map from food-code digits to taxonomy labels. Level.code is the code prefix, Main.food.description its label and n_digits the level. source records provenance: usda_fndds for labels taken from the official USDA scheme, food_tree_extension for the finer levels, which come from the food-tree method of Johnson et al. and have no USDA counterpart. usda_fndds_description gives the official wording where one exists, and used_in_F1d_tree flags the labels the published tree uses. final_table_for_writing_out_to_newick.csv (622 rows, one per food) — the full food taxonomy: FoodID and description, the seven taxonomy levels L1 to L7, the newickstring and newickstring_name paths, and a semicolon-joined taxonomy string. output_food_tree_datatree_name.newick — the food taxonomy tree in Newick format with descriptive leaf labels; this is the file the renderer consumes. output_food_tree_datatree.newick — the same tree with food-code leaf labels. food_group_color_key_final.csv (9 rows) — the nine major food groups: numeric code fgrp1, full description fdesc, the column name used in the analysis tables fg1_name, the plotting color as a hex value, and a shortname. annotation.base.txt — the GraPhlAn configuration header used when rendering the circular tree. United States Department of Agriculture reference files, redistributed unmodified for convenience. These are United States government works in the public domain; the authoritative copies remain with the USDA. Each workbook carries its own variable dictionary as a second worksheet. FPED_1516.xls, FPED_1720.xls — Food Patterns Equivalents Database, 2015–2016 and 2017–2020. One row per food code giving its content in food-pattern equivalents (cup equivalents of fruit and vegetable subgroups, ounce equivalents of protein foods, teaspoon equivalents of added sugars, and so on). Used to split total sugars into added and other sugars. The Variable_labels sheet defines all 39 columns with units. 2019-2020 FNDDS At A Glance - FNDDS Nutrient Values.xlsx — Food and Nutrient Database for Dietary Studies nutrient values per 100 g of edible portion, used for the food codes absent from the 2015–2016 release supplied as Data S5. The Variable Descriptions sheet defines all 69 columns. Code availability. The analysis code that reproduces the public figures of the manuscript from these data is archived at https://doi.org/10.5281/zenodo.21290618 (concept DOI, always resolves to the latest release) and developed at https://github.com/Anqi-Dai/Nutrition_microbiome_reproducibility. Additional File: Filled-out STORMS checklist This file is a filled-out STORMS checklist for the manuscript. It is version 1.03, downloaded from 10.5281/zenodo.5703116. The STORMS checklist is a standardized checklist for microbiome studies, published in the journal Nature Medicine (https://www.nature.com/articles/s41591-021-01552-x).

提供机构:
Zenodo
创建时间:
2026-07-31
二维码
社区交流群
二维码
科研交流群
商业服务