Supporting datasets for Comparison of Multi-locus Sequence Typing software for next generation sequencing data
收藏资源简介:
To test the accuracy of MLST applications, we have constructed two datasets of simulated reads in FASTQ format. The first has perfect reads over the MLST genes, plus a flanking region (based on the Salmonella Typhi CT18 reference) in varying levels of coverage from 1 to 30. This allows for us to see at what point each software application can accurately detect an allele.The second dataset is similar to the first, but contains 2 Salmonella samples Salmonella Typhi CT18 and Salmonella Weltevreden, with the samples mixed in varying ratios. This allows us to see at what point software applications detect that there is a mixed allele/contamination.
为验证多位点序列分型(MLST)工具的应用准确性,本研究构建了两份FASTQ格式的模拟测序读段数据集。第一份数据集包含覆盖MLST基因区域及侧翼序列的无错测序读段,其测序覆盖度基于伤寒沙门氏菌CT18参考基因组设置,覆盖度范围为1至30倍且梯度可变。借此可观测不同测序覆盖度下,各软件工具准确识别等位基因的临界阈值。第二份数据集与第一份结构相似,但包含两份沙门氏菌样本——伤寒沙门氏菌CT18与韦太夫雷登沙门氏菌,并以不同比例将两份样本混合。借此可观测软件工具识别混合等位基因或样本污染的临界阈值。



