遇见数据集

LIMNOS Metagenomes Gene Catalog

收藏
Zenodo2026-07-24 更新2026-08-01 收录
官方服务:

资源简介:

This repository contains the LIMNOS Gene catalog. A full description of the metholodogy used to create this catalog can be found below. This repository contains five files LAKE_HPG_MAG_9590.faa.gz, LAKE_HPG_MAG_9590.fasta.gz, LAKE_HPG_MAG_9590.clstr.gz, LAKE_HPG_ALL_GENES.kegg.tsv.gz, LAKE_HPG_MAG_9590-profile_insertcounts.lengthnorm.profile.cellab.profile.gz The gene catalog uses the term gene to refer to redundant gene sequences identified in prokaryotic MAGs (https://zenodo.org/records/10533926) and gene_clusters to refer to non-redundant sequences after clustering. This gene catalog contains approximately 180 Mio genes, forming ~ 18 Mio non-redundant gene clusters, recovered from ~76K MAGs. Of these 18 Mio non-redundant gene_clusters we were able to assign functional annotation to around 5 Mio using the KEGG database and DIAMOND blast tool. (see methods below). LAKE_HPG_MAG_9590.faa.gz and LAKE_HPG_MAG_9590.fasta, contain the non-redundant gene catalog in amino acid and nucleotide fasta format, respectively. LAKE_HPG_MAG_9590.clstr.gz is a mapping file linking genes from invidual MAGs to the non-redundant gene clusters. LAKE_HPG_ALL_GENES.kegg.tsv.gz contains the functions assigned to the individual genes forming the redundant gene catalog. LAKE_HPG_MAG_9590-profile_insertcounts.lengthnorm.profile.cellab.profile.gz is a table containing the length and cell abundance normalised abundances of each of the individual gene_clusters in all 1097 samples comprising the LIMNOS dataset. Note: Abundances are scaled by 1000 to ensure sharabilty. We recommend dividing each value in the table by 1000 before using for ecological exploration, to avoid surprising results. Data in this file can be interpreted as copies of the gene_cluster per cell (or genome) in the population. For example a value of 0.01 suggests either 1% of prokaryotic genomes contain this gene, or that 0.5% of prokaryotic genomes contain this gene (assuming it is present twice per genome). See references below for more details. Sampling and DNA extraction Water samples were collected at various time-points between 2010 and 2020 from four lakes in Northeastern Germany. Regular sampling approximates monthly sampling, with the exception of ice-covered periods when where sampling was impossible or prohibited, and during 2020 when where biweekly samples were collected from Lake Stechlin. Epilimnion (STE samples (n = 804), represented the majority of samples taken from the four lakes, and constitute represent integrated samples taken from a set of discrete depths. Outside of the seasonally stratified periods “epilimnetic epilimnic” samples were taken from a defined depth of 5 m. In Lake Fuchskuhle, epilimnetic epilimnic samples were taken from a depth of 1 m (total depth 4.5 m) independent of lake stratification stratified periods (with a strong oxycline occurring around 2 m during stratification). Amongst the remaining samples, a large number represent samples from the hypolimnion (STH, 40 m) and deep-chlorophyll maxima (DCM; 10-15 m) of Lake Stechlin, or depth discrete samples (0-4 m) from Lake Fuchskuhle. Water samples (300-1000 ml) were transported in cool-boxes, to the laboratory after no more than 4 hours, sequentially filtered through 5 µm and 0.22 µm Millipore cellulose nitrate membranes. Membrane filters were stored at -20oC for a maximum period of 10 years. Total DNA was extracted from cellulose-nitrate membranes using a modified phenol-chloroform protocol (1). Prior to DNA sequencing purified DNA was stored in 1 mM Tris-HCl, pH 8.0 at -20oC for no longer than 6 months. Metagenome sequencing Metagenome sequencing was performed at the Ramaciotti Centre for Genomics, Sydney, Australia. DNA extracts were transferred into 96 well plates, desiccated under vacuum at ambient temperature and packaged with ice-packs in polystyrene. Receipt of samples by the sequencing facility was made within 4 days of shipping, with samples stored at -20oC on arrival. Metagenome sequencing libraries were prepared using Illumina DNA Prep kits according to manufacturer’s specifications. Paired end 2x150 bp sequencing was undertaken using the NovaSeq 6000 system and the using NovaSeq 6000 S4 Reagent Kit v1.5. To the reduced the likelihood of batch effects, samples were randomized and sequenced across four full NovaSeq runs. Raw paired-end sequence files were submitted to the European Nucleotide Archive (PRJEB47226, SAMEA9560359- SAMEA9561456). In addition, raw metagenome reads, from 267 freshwater lakes were downloaded from a previous study (PRJEB38681, (2)). Metagenome assembly and binning Metagenomic (n=1,363) sequencing datasets were processed as previously described in (3). Briefly, BBMap (v.38.71) was used to quality control sequencing reads from all samples by removing adapters from the reads, removing reads that mapped to quality control sequences (PhiX genome) and discarding low-quality reads (trimq=14, maq=20, maxns=1, and minlength=45). Quality-controlled reads were merged using bbmerge.sh with a minimum overlap of 16 bases, resulting in merged, unmerged paired, and single reads. The reads from metagenomic samples were assembled into scaffolded contigs (hereafter scaffolds) using the SPAdes assembler (v3.15.2) (4) in metagenomic mode. Assemblies were submitted to the ENA (PRJEB47226). Scaffolds were length-filtered (≥ 1,000 bp) and quality-controlled reads from each metagenomic sample were mapped against the scaffolds of each sample. Mapping was performed using BWA (v0.7.17-r1188; -a) (5). Alignments were filtered to be at least 45 bp in length, with an identity of ≥ 97% and a coverage of ≥ 80% of the read sequence. The resulting BAM files were processed using the jgi_summarize_bam_contig_depths script of MetaBAT2 (v2.12.1) (6) to compute within- and between-sample coverages for each scaffold. The scaffolds were binned by running MetaBAT2 on all samples individually. The resulting BAM files were processed using the jgi_summarize_bam_contig_depths script of MetaBAT2 (v.2.12.1, (6) to provide within- and between-sample coverages for each scaffold. The scaffolds were finally binned by running MetaBAT 2 on all samples individually with parameters --minContig 2000 and --maxEdges 500 for increased sensitivity. The quality of each metagenomic bin and external genome was evaluated using both the ‘lineage workflow’ of CheckM (v.1.1.3, (7)) and Anvi’o (v.7.1, (8)). Metagenomic bins were retained for downstream analyses if either CheckM or Anvi’o reported a completeness/completion (cpl) of ≥50% and a contamination/redundancy (ctn) of ≤10%. Prokaryotic genomes were taxonomically annotated using GTDB-Tk (v.2.1.0, (9)) with the default parameters against the GTDB R202 release (10). The full MAG catalog is available online (10.5281/zenodo.10533926) Gene catalog Gene sequences were predicted from prokaryotic MAGs using Prokka (v1.14.5, (11)) with default parameters. Genes from prokaryotic MAGs were subsequently clustered at 95% identity, keeping the longest sequence as representative using CD-HIT (v4.8.1) with the parameters -c 0.95 -M 0 -G 0 -aS 0.9 -g 1 -r 0 -d 0 -b 1000. Representative gene sequences were aligned against the KEGG database (release April.2022) using DIAMOND (v2.0.15) (12) and filtered to have a minimum query and subject coverage of 70% and requiring a bitScore of at least 50% of the maximum expected bitScore (reference against itself). The 1,097 LIMNOS metagenomes were then mapped to the ~18 million cluster representatives with BWA (v0.7.17-r1188; -a) (5) and the resulting BAM files were filtered to retain only alignments with a percentage identity of ≥95% and ≥45 bases aligned. Length-normalised gene abundance was calculated by first counting inserts from best unique alignments and then, for ambiguously mapped inserts, adding fractional counts to the respective target genes in proportion to their unique insert abundances and dividing the total insert counts by the length of the respective gene. Gene-length normalised read abundances were further converted into per-cell gene copy numbers dividing them by the median abundance of single-copy marker genes copies in each sample (3). The full gene catalog is available online (10.5281/zenodo.10533926). 1. D. Ionescu, M. Bizic, R. Karnatak, C. L. Musseau, G. Onandia, M. Kasada, S. A. Berger, J. C. Nejstgaard, M. Ryo, G. Lischeid, M. O. Gessner, S. Wollrab, H. ‐P. Grossart, From microbes to mammals: Pond biodiversity homogenization across different land‐use types in an agricultural landscape. Ecological Monographs 92, e1523 (2022). 2. M. Buck, S. L. Garcia, L. Fernandez, G. Martin, G. A. Martinez-Rodriguez, J. Saarenheimo, J. Zopfi, S. Bertilsson, S. Peura, Comprehensive dataset of shotgun metagenomes from oxygen stratified freshwater lakes and ponds. Sci Data 8, 131 (2021). 3. G. Salazar, L. Paoli, A. Alberti, J. Huerta-Cepas, H.-J. Ruscheweyh, M. Cuenca, C. M. Field, L. P. Coelho, C. Cruaud, S. Engelen, A. C. Gregory, K. Labadie, C. Marec, E. Pelletier, M. Royo-Llonch, S. Roux, P. Sánchez, H. Uehara, A. A. Zayed, G. Zeller, M. Carmichael, C. Dimier, J. Ferland, S. Kandels, M. Picheral, S. Pisarev, J. Poulain, S. G. Acinas, M. Babin, P. Bork, E. Boss, C. Bowler, G. Cochrane, C. De Vargas, M. Follows, G. Gorsky, N. Grimsley, L. Guidi, P. Hingamp, D. Iudicone, O. Jaillon, S. Kandels-Lewis, L. Karp-Boss, E. Karsenti, F. Not, H. Ogata, S. Pesant, N. Poulton, J. Raes, C. Sardet, S. Speich, L. Stemmann, M. B. Sullivan, S. Sunagawa, P. Wincker, S. G. Acinas, M. Babin, P. Bork, C. Bowler, C. De Vargas, L. Guidi, P. Hingamp, D. Iudicone, L. Karp-Boss, E. Karsenti, H. Ogata, S. Pesant, S. Speich, M. B. Sullivan, P. Wincker, S. Sunagawa, Gene Expression Changes and Community Turnover Differentially Shape the Global Ocean Metatranscriptome. Cell 179, 1068-1083.e21 (2019). 4. S. Nurk, D. Meleshko, A. Korobeynikov, P. A. Pevzner, metaSPAdes: a new versatile metagenomic assembler. Genome Res. 27, 824–834 (2017). 5. H. Li, R. Durbin, Fast and accurate short read alignment with Burrows–Wheeler transform. Bioinformatics 25, 1754–1760 (2009). 6. D. D. Kang, F. Li, E. Kirton, A. Thomas, R. Egan, H. An, Z. Wang, MetaBAT 2: an adaptive binning algorithm for robust and efficient genome reconstruction from metagenome assemblies. PeerJ 7, e7359 (2019). 7. D. H. Parks, M. Imelfort, C. T. Skennerton, P. Hugenholtz, G. W. Tyson, CheckM: assessing the quality of microbial genomes recovered from isolates, single cells, and metagenomes. Genome Res. 25, 1043–1055 (2015). 8. A. M. Eren, Ö. C. Esen, C. Quince, J. H. Vineis, H. G. Morrison, M. L. Sogin, T. O. Delmont, Anvi’o: an advanced analysis and visualization platform for ‘omics data. PeerJ 3, e1319 (2015). 9. P.-A. Chaumeil, A. J. Mussig, P. Hugenholtz, D. H. Parks, GTDB-Tk: a toolkit to classify genomes with the Genome Taxonomy Database. Bioinformatics 36, 1925–1927 (2020). 10. D. H. Parks, M. Chuvochina, D. W. Waite, C. Rinke, A. Skarshewski, P.-A. Chaumeil, P. Hugenholtz, A standardized bacterial taxonomy based on genome phylogeny substantially revises the tree of life. Nat Biotechnol 36, 996–1004 (2018). 11. T. Seemann, Prokka: rapid prokaryotic genome annotation. Bioinformatics 30, 2068–2069 (2014). 12. B. Buchfink, C. Xie, D. H. Huson, Fast and sensitive protein alignment using DIAMOND. Nat Methods 12, 59–60 (2015).

提供机构:
Zenodo
创建时间:
2026-07-24
二维码
社区交流群
二维码
科研交流群
商业服务