Simulated metagenomes with quality and abundance distributions derived from real samples
收藏资源简介:
Species abundances and quality values were derived from the following list of samples: <pre><code>SAMEA2466896 SAMEA2466916 SAMEA2466952 SAMEA2466953 SAMEA2466965 SAMEA2466996 SAMEA2467015 SAMEA2467039 SAMEA2621010 SAMEA2621033 SAMEA2621107 SAMEA2621155 SAMEA2621229 SAMEA2621247 SAMEA2621300 SAMEA2622357 </code></pre> Reference abundances were generated using mOTUs profiler.<br> Metagenomes were simulated with cMESSi using proGenomes' representative contigs for species and the aforementioned abundances. In cases where a <em>ref_mOTU_v2</em> corresponded to more than one genome, the abundance of said <em>ref_mOTU</em> was distributed equally over all genomes.<br> GFF location files were produced using location information generated by cMESSi.<br> Truth vectors were obtained by intersecting coordinates of simulated reads with coordinates of eggNOG orthologous groups (OG at NOG level) as predicted by eggNOG-mapper. A read overlapping multiple genes is considered for each gene. If a gene possesses multiple NOG annotations, each annotation gets assigned the total number of overlapping reads. If you use this dataset, please cite: NG-meta-profiler: fast processing of metagenomes using NGLess, a domain-specific language



