遇见数据集

SFB genomes and annotations

收藏
Zenodo2019-06-18 更新2026-04-07 收录
数据链接:
官方服务:

资源简介:

This dataset contains sequence files for a Metagenome Assembled Genome (MAG) from human metagenomes, as well as 5 SFB reference genomes: GCF_000270205 Candidatus Arthromitus sp. SFB-mouse-Japan GCF_000283555 Candidatus Arthromitus sp. SFB-rat-Yit GCF_000284435 Candidatus Arthromitus sp. SFB-mouse-Yit GCF_000709435 Candidatus Arthromitus sp. SFB-mouse-NL GCF_001655775 Candidatus Arthromitus sp. SFB-turkey isolate UMNCA01 The dataset consists of 8 gzipped tar archives. Here's brief summary of their contents: <strong>sfb_abundance</strong>: Counts of mapped reads and normalized counts for each contig in 825 samples (see <strong>sfb_map</strong>) Files named 'raw_counts' are number of reads assigned to each contig while files named 'tpm' are counts normalized to Transcripts Per Million. The 'percontig' files show numbers per contig while raw_counts.tab and tpm.tab files have counts summed for each genome. <strong>sfb_abundance.cds</strong>: Counts of mapped reads and normalized counts as above but only for reads mapping to protein-coding regions. <strong>sfb_annotations</strong>: Annotation files, from running the prokka pipeline on the genomes and subsequently eggnog-mapper, pfam_scan and dbCAN. <strong>sfb_checkm</strong>: Results from running 'checkm lineage_wf' on the genomes. <strong>sfb_collated</strong>: Collated counts of annotations in each genome. <strong>sfb_fastani</strong>: Results from running fastANI on the genomes, with subsequent clustering of genomes based on 75% overlap and 95% ANI. <strong>sfb_gtdb</strong>: Results from the 'gtdbtk classify_wf' on the genomes. This shows how the genomes are classified against the Genome Taxonomy Database (release86). <strong>sfb_gtdb_denovo</strong>: Phylogeny as created using the following command on the genomes. <pre><code class="language-bash">gtdbtk de_novo_wf --bac120_ms --outgroup_taxon p__Patescibacteria -x .fna --cpus 20 --rnd_seed 123</code></pre> <strong>sfb_map</strong>: Results from mapping reads from 825 samples to the 6 genomes. Reads were aligned using bowtie2 with '--very-sensitive --no-unal' settings and '--score-min C,0,0' to only report reads aligning without mismatches.Output was sorted by position and duplicates removed using MarkDuplicates of the picard tools suite. The archive contains a single merged bam file ('sfb.bam') where each sample has been assigned a ReadGroup inferred from its file name. Note that this mapping step was performed to investigate the presence of the SFB MAG in other metagenomes and was not part of the actual binning step. <strong>SFB.unoise.vsearch.tsv: </strong>Count table of amplified 16S sequence variants with one sample per column and one Amplicon Sequence Variant (ASV) per row. The sixth column shows the assigned taxonomy, and the seventh, the sequence. Total DNA was amplified with the universal bacterial 16S primer pair 341f-805r. Primer sequences and low quality bases were removed from the raw reads with Cutadapt. ASVs were picked using Unoise3 with standard parameters. Taxonomy was assigned by the SINA classifier, based on the SILVA database v132.

创建时间:
2019-06-18
二维码
社区交流群
二维码
科研交流群
商业服务