Comprehensive Analysis to Improve the Validation Rate for Single Nucleotide Variants Detected by Next-Generation Sequencing
收藏资源简介:
Next-generation sequencing (NGS) has enabled the high-throughput discovery of germline and somatic mutations. However, NGS-based variant detection is still prone to errors, resulting in inaccurate variant calls. Here, we categorized the variants detected by NGS according to total read depth (TD) and SNP quality (SNPQ), and performed Sanger sequencing with 348 selected non-synonymous single nucleotide variants (SNVs) for validation. Using the SAMtools and GATK algorithms, the validation rate was positively correlated with SNPQ but showed no correlation with TD. In addition, common variants called by both programs had a higher validation rate than caller-specific variants. We further examined several parameters to improve the validation rate, and found that strand bias (SB) was a key parameter. SB in NGS data showed a strong difference between the variants passing validation and those that failed validation, showing a validation rate of more than 92% (filtering cutoff value: alternate allele forward [AF]≥20 and AF
下一代测序(Next-generation sequencing, NGS)技术已实现生殖系与体细胞突变的高通量发掘。然而,基于NGS的变异检测仍易出现误差,导致变异分型结果不准确。本研究依据总读长深度(total read depth, TD)与单核苷酸多态性质量(SNP quality, SNPQ)对NGS检出的变异进行分类,并选取348个非同义单核苷酸变异(non-synonymous single nucleotide variants, SNVs)开展桑格测序(Sanger sequencing)验证。借助SAMtools与GATK算法分析后发现,验证率与SNPQ呈显著正相关,但与TD无明显关联。此外,两款工具共同检出的常见变异的验证率高于仅单一工具检出的变异。本研究进一步考察了多项可提升验证率的参数,发现链偏好性(strand bias, SB)为关键参数。NGS数据中的链偏好性在验证通过与验证失败的变异间存在显著差异,其验证率可达92%以上(筛选临界值:替代等位基因正向链比例[AF]≥20且AF )



