The metagenome of the naked mole-rat
收藏资源简介:
Naked mole-rat metagenome and gene catalog Contents of the dataset Output related to the 16S rRNA gene sequencing (001_16s_related_files): Directory with QIIME2 output (QZA format): qiime2.tar.gz. Below is the description of each file: Raw imported sequences from all datasets, including NMR and mouse data: 20260209_16_33_25-pooled-single-end-demux.qza. The QZA file includes metadata that categorises samples by host Sequences with primers trimmed by cutadapt: 20260209_16_33_25-pooled-single-end-trimmed.qza Feature table generated by DADA2 pipeline: 20260209_16_33_25-pooled-single-trimmed-dada2-table-234.qza Representative sequences generated by DADA2 pipeline: 20260209_16_33_25-pooled-single-trimmed-dada2-rep-seqs-234.qza Denoising statistics generated by DADA2: 20260209_16_33_25-pooled-single-trimmed-dada2-stats-234.qza Representative sequences filtered by length (>199 nt): 20260209_16_33_25-pooled-single-trimmed-dada2-rep-seqs-234-filtered.qza Feature table filtered by representative sequences (only sequences longer than 199 nt are retained): 20260209_16_33_25-pooled-single-trimmed-dada2-table-234-filtered.qza Taxonomic classification of filtered representative sequences: 20260209_16_33_25-pooled-single-trimmed-dada2-234-filtered-taxonomy.qza Alignment of filtered representative sequences by MAFFT: 20260209_16_33_25-pooled-single-trimmed-dada2-rep-seqs-234-filtered-aligned.qza Masked alignment of filtered representative sequences by MAFFT: 20260209_16_33_25-pooled-single-trimmed-dada2-rep-seqs-234-filtered-aligned-masked.qza Rooted tree built by MAFFT: 20260209_16_33_25-pooled-single-trimmed-dada2-234-filtered-rooted-tree.qza Unrooted tree built by MAFFT: 20260209_16_33_25-pooled-single-trimmed-dada2-234-filtered-unrooted-tree.qza Taxonomic classification of filtered representative sequences by SILVA database: 20260209_16_33_25-pooled-single-trimmed-dada2-234-filtered-taxonomy.qza Directories with MaAsLin2 output: Comparison of naked mole-rats with other rodents (genus level): 20260213_12_10_45-NMR-DMR-B6mouse-MSMmouse-FVBNmouse-spalax-pvo-hare-rabbit-Genus-host-234-ref-NMR Comparison of naked mole-rat sexes (with relations as random_effects parameter) (ASV level): 20260215_19_26_58-NMR-OTU-sex-234-ref-female-with-relations Comparison of naked mole-rat age groups (with relations as random_effects parameter) (ASV level): 20260213_14_38_30-NMR-OTU-age-234-ref-agegroup0_10-with-relations Tables of differentially abundant taxa from MaAsLin2 (TSV): Comparison of naked mole-rats with other rodents (genus level): 20260213_12_20_50-maaslin2-NMR-DMR-B6mouse-MSMmouse-FVBNmouse-spalax-pvo-hare-rabbit-Genus-host-234-ref-NMR-signif.tsv Comparison of naked mole-rat sexes (ASV level): 20260215_19_30_37-maaslin2-NMR-OTU-sex-234-ref-female-signif.tsv.tsv Comparison of naked mole-rat age groups (with relationships as random_effects parameter) (ASV level): 20260213_14_39_57-maaslin2-NMR-OTU-age-234-ref-agegroup0_10-signif.tsv Tables of differentially abundant features from ALDEx2 (features that have CI not overlapping zero and a high effect size) (TSV): Comparison of naked mole-rats with other rodents (genus level): 20260213_13_20_05-aldex2-NMR-DMR-B6mouse-MSMmouse-FVBNmouse-spalax-pvo-hare-rabbit-Genus-host-234-ref-NMR-signif.tsv Table of differentially abundant taxa from ANCOM-BC: Comparison of naked mole-rats with other rodents (genus level): 20260213_13_41_58-ancombc-NMR-DMR-B6mouse-MSMmouse-FVBNmouse-spalax-pvo-hare-rabbit-Genus-host-234-ref-NMR-signif.tsv Filtered table of differentially abundant taxa from MaAsLin2 (negative coef in all hosts except reference - NMR) (TSV): 20260213_12_20_51-maaslin.signif.decreased-NMR-DMR-B6mouse-MSMmouse-FVBNmouse-spalax-pvo-hare-rabbit-Genus-host-234-ref-NMR-signif.tsv Filtered table of differentially abundant taxa from ALDEx2 (negative values in all hosts except reference - NMR) (TSV): 20260213_13_19_00-aldex.neg.effect-NMR-DMR-B6mouse-MSMmouse-FVBNmouse-spalax-pvo-hare-rabbit-Genus-host-234-ref-NMR-signif.tsv Filtered table of differentially abundant taxa from ANCOM-BC (negative effect in all hosts except reference - NMR) (TSV): 20260213_13_42_00-ancombc.signif.decreased-NMR-DMR-B6mouse-MSMmouse-FVBNmouse-spalax-pvo-hare-rabbit-Genus-host-234-ref-NMR-signif.tsv Table of common differentially abundant taxa between MaAsLin2 and ANCOM-BC (decreased in all hosts except reference - NMR) (TSV): 20260213_13_54_58-significant-features-NMR-DMR-B6mouse-MSMmouse-FVBNmouse-spalax-pvo-hare-rabbit-Genus-host-234-ref-NMR-signif.tsv Output related to the whole metagenome sequencing (Kraken2/Bracken pipeline) (002_wms_related_files): Taxonomy table from Bracken (all samples combined) (TSV): 20240515_10_04_04_combined_report.tsv MaAsLin2 output (directory): comparison of NMR ages with relations as random_effects parameter (Kraken2): 20240621_18_31_01-NMR-nonfiltered-Species-age-agegroup0_10-agegroup10_16-ref-agegroup0_10 Output related to MAG assembly (003_mag_assembly_related_files): Directory with contigs assembled by MEGAHIT: megahit.tar.gz Directory with MAGs assembled by MetaBAT2: metabat2.tar.gz. Includes depth files Table of contigs assembled by MEGAHIT in each sample, with length and GC% calculated by SeqKit: 20250712_18_07_37_megahit_combined_stats.tsv Table of contigs assembled by MEGAHIT and binned by MetaBAT2 in each sample, with length GC% calculated by SeqKit: 20250712_18_07_37_metabat2_combined_stats.tsv CheckM2 output with reports: checkm2.tar.gz. Each sample has its own subdirectory with the contents as described in the documentation. CheckM2 report on all MAGs: 20250619_05_47_09_quality_reports_merged.tsv CheckM2 report on high-quality MAGs: 20250619_05_47_09_high_quality_mags.tsv Representative MAG information from dRep at species level (secondary clustering threshold of 95% ANI) (drep95.tar.gz): data/ : Clustering_files, fastANI_files, and MASH_files (as described in the documentation) data_tables/ : CSV files from dRep (as described in the documentation). Contains the representative bin IDs (sample and bin ID) (Wdb.csv). dereplicated_genomes/ : FASTA files with representative MAGs figures/ : figures generated by dRep log/ : log files Representative MAG information from dRep at strain level (secondary clustering threshold of 99% ANI) (drep99.tar.gz): same structure as drep95.tar.gz Table of contigs found in high-quality dereplicated bins (dRep), with length GC% calculated by SeqKit: 20250712_18_07_37_sa_95perc_combined_stats.tsv Intermediate output from GTDB-tk on species bins from drep95.tar.gz: 20250705_19_24_02_sa_95perc_GTDBtk.tar.gz. Contains the align/, classify/, and identify/ folders, and log files. Intermediate output from GTDB-tk on strain bins from drep99.tar.gz: 20250705_19_24_02_sa_99perc_GTDBtk.tar.gz. Contains the align/, classify/, and identify/ folders, and log files. Final taxonomic classification table of bins by GTDB-tk (from drep95.tar.gz): 20250705_19_24_02_sa_95perc_classification_combined.tsv Final taxonomic classification table of bins by GTDB-tk (from drep99.tar.gz): 20250705_19_24_02_sa_99perc_classification_combined.tsv BLASTN classification of contigs (SILVA SSU Ref NR99 reference database) (blastn.tar.gz): *_SILVA_results.tsv: Raw BLASTN output for each sample *_SILVA_results_filtered.tsv: Filtered BLASTN output (sequence identity > 90% and alignment length > 500 bp) *_results_for_taxonomy.tsv: top hits (by bitscore) from *_SILVA_results_filtered.tsv, but only containing qseqid, sseqid, pident, length, bitscore *_blast_with_taxonomy.tsv: top hits matched with SILVA taxonomy by IDs (sseqid) 20250701_05_24_10_blastn_all_samples.tsv: Final combined BLASTN output from all samples Relative abundances of all MAGs calculated by CoverM: 20250803_05_07_53_all_samples_bins_coverm.tsv Relative abundances of representative species MAGs calculated by CoverM: 20250803_05_07_53_all_samples_sa_95perc_coverm.tsv Relative abundances of all contigs calculated by CoverM: 20250803_05_07_53_all_samples_megahit_coverm.tsv Table with taxonomic classification of each contig with BLASTN and GTDB-tk, contig length, mapping to the corresponding bin, and indication whether the contig belongs to a representative bin, or not: seqkit-with_taxonomy.tsv TPM of contigs classified by BLASTN, calculated with CoverM: blastn-coverm-megahit.tsv Relative abundance of all MAGs classified by GTDB-tk, calculated with CoverM: gtdbtk-coverm-metabat2.tsv Relative abundance of representative species MAGs calculated with CoverM: gtdbtk-coverm-drep.tsv Directory with output from PROKKA: prokka.tar.gz. Includes all the usual files (FAA, FASTA, GTF, etc), unprocessed FASTA file with all predicted proteins: 20250626_22_11_43_all_proteins.faa Mapping of each contig to predicted gene (locus_tag): 20250626_22_11_43_all_contig_gene_maps.tsv Clustering output from MMseqs2 easy-cluster: mmseqs2.tar.gz: 20250626_22_11_43_mmseqs_easy_cluster_report.txt: log file 20250626_22_11_43_prokka_nr_prot_ids.txt: representative CDS IDs 20250626_22_11_43_prokka_nr_prot_rep_seq.faa: FASTA file with representative CDS that were used in KofamScan and dbCAN Mapping of each locus_tag from PROKKA (predicted CDS) to locus_tag_cluster from MMseqs2 (representative CDS): 20250626_22_11_43_prokka_nr_prot_cluster.tsv Classification of representative CDS with KofamScan (unfiltered): 20250718_09_21_41_nr_prot_kofam_scan.txt Classification of representative CDS with KofamScan (top hits by threshold and score): 20250718_09_21_41_nr_prot_kofam_scan_top_hits.txt Classification of representative CDS with dbCAN: 20250626_22_11_43_nr_dbcan_overview.txt All KO definitions identified in 20250718_09_21_41_nr_prot_kofam_scan.txt: ko-defs.tsv Representative gene lengths: representative-gene_lengths.tsv Annotation of each predicted CDS (locus_tag) with mapping to the representative ID (locus_tag_cluster), corresponding sample and contig, CDS length, TPM, and annotation by dbCAN (CAZ, CAZ combination, CAZ subclass, and CAZ class) and KofamScan: gene-annotation-df.tsv Classification of CAZymes with BLASTP (all hits) on dbCAN reference database: 20250802_12_29_03_blastp_cazy_db_cazymes_results.tsv Classification of CAZymes with BLASTP (only top hits) on dbCAN reference database: 20250802_12_29_03_blastp_cazy_db_cazymes_top_hits_with_taxa.tsv Counts of mapped reads on each gene with htseq-count (htseq-count.tar.gz): 20250710_19_36_58_*_prokka.count: raw output for each sample 20250710_19_36_58_*_prokka_filtered.count: only rows that start with "ID=", with "ID=" removed All 20250710_19_36_58_*_prokka_filtered.count files combined into one: 20250710_19_36_58_all_prokka_counts.tsv



