RefSeq Contamination
收藏资源简介:
Contamination prediction for RefSeq (July 2019) by conterminator (https://github.com/martin-steinegger/conterminator).<br>The predictions are TSV formated. The following is the column defintion:<br><br>1.) Numeric identifier<br>2.) Contaminated identifier<br>3.) Kingdom (0: Bacteria&Archaea, 1: Fungi, 2: Metazoa, 3: Viridiplantae, 4: Other Eukaryotes)<br>4.) Species name<br>5.) Alignment start<br>6.) Alignment end<br>7.) Corrected contig length<br>8.) Identifier of the longest contaminating sequence<br>9.) Kingdom of the longest contaminating sequence<br>10.) Species name of the longest contaminating sequence<br>11.) Length of the longest contaminating sequence<br>12.) Count how often sequences from contaminating kingdom align<br><br>Contaminated identifiers can occur multiple times if multiple alignments were detected.
本数据集为使用conterminator工具(https://github.com/martin-steinegger/conterminator)针对2019年7月版参考序列数据库(RefSeq)完成的污染预测结果。 预测结果采用制表符分隔值(TSV)格式存储,各字段定义如下: 1. 数值型标识符 2. 污染序列标识符 3. 分类界(0:细菌与古菌,1:真菌,2:后生动物,3:绿色植物,4:其他真核生物) 4. 物种名称 5. 比对起始位点 6. 比对终止位点 7. 校正后的重叠群长度 8. 最长污染序列的标识符 9. 最长污染序列的分类界 10. 最长污染序列的物种名称 11. 最长污染序列的长度 12. 污染分类界内序列的比对频次统计 若检测到多处比对,同一污染序列标识符可能会重复出现。



