GORG Dark - Reference Contaminant Dataset
收藏资源简介:
A reference database of contaminant sequences that were screened during the creation of the GORG-Dark dataset. Context: GORG-Dark is a set of 9,698 single amplified genomes (SAGs) generated from individual prokaryotes picked from the aphotic/deep ocean. Purpose: This contaminant dataset was used in the bioinformatic screening of contaminant reads and contigs during SAG assembly. Dataset creation: This dataset is a combination of both 'in-house' contaminants and public contaminant genomes. The in-house sequences are lab/reagent contaminants detected at the Single Cell Genomics Center by amplifying SAGs from empty wells (i.e. 384 well plates containing only buffer, with no deliberately introduced DNA). Dominant in-house contaminants are DNA from humans, Acinetobacter, Delftia, Achromobacter, Bradyrhizobium, Stenotrophiomonas, and other opportunisitic gram-negative bacteria. The public sequences consist of the human assembly GRCh38 (which constitutes ~45% of the total database), as well as the mouse genome mm10 (~43% of the total) as well as public assemblies of common skin microbes Bradyrhizobium, Ralstonia, Malassezia, and Cutibacterium. All input sequences were masked to hide low-complexity sequences prior to combining into the final fasta file GRCh38_AG665_mm10.fa. This fasta was then indexed for Burrows Wheeler Alignment searching (by BWA, v0.7.17, using the command 'bwa index GRCh38_AG665_mm10.fa'). The resulting BWA index files (*amb, *ann, *bwt, *pac, *sa) are included here. All files must be unzipped with the command unzip, prior to use.



