Draft assembly of Nardus stricta
收藏资源简介:
We assembled the species using a pre-release of the EBP-Nor genome assembly pipeline (https://github.com/ebp-nor/GenomeAssembly). KMC (Kokot et al., 2017) was used to count k-mers of size 32 in the PacBio HiFi reads, excluding k-mers occurring more than 10,000 times. GenomeScope (Ranallo-Benavidez et al., 2020) was run as part of the pipeline on the k-mer histogram output from KMC and was included in the methods for completeness. Ploidy level was investigated using Smudgeplot (Ranallo-Benavidez et al., 2020). HiFiAdapterFilt (Sim et al., 2022) was applied on the HiFi reads to remove possible remnant PacBio adapter sequences. The filtered HiFi reads were assembled using Hifiasm (Cheng et al., 2021) with Hi-C integration resulting in a pair of haplotype-resolved assemblies, pseudo-haplotype one (hap1) and pseudo-haplotype two (hap2). Unique k-mers in each assembly/pseudo-haplotype were identified using meryl (Rhie et al., 2020) and used to create two sets of Hi-C reads, one without any k-mers occurring uniquely in hap1 and the other without k-mers occurring uniquely in hap2. K-mer filtered Hi-C reads were aligned to each scaffolded assembly using BWA-MEM (Li, 2013) with -5SPM options. The alignments were sorted based on name using samtools (Li et al., 2009) before applying samtools fixmate to remove unmapped reads and secondary alignments and to add mate score, and samtools markdup to remove duplicates. The resulting BAM files were used to scaffold the two assemblies using YaHS (Zhou et al., 2023) with default options. FCS-GX (Astashyn et al., 2024) was used to search for putative contamination and contaminated sequences were removed. lpNarStri2.hap1.decon.fasta.gz - genome assembly of Nardus stricta (hap1) lpNarStri2.hap2.decon.fasta.gz - genome annotation of Nardus stricta (hap2)
本研究采用EBP-Nor基因组组装流程的预发布版本完成物种基因组组装,该流程的开源地址为https://github.com/ebp-nor/GenomeAssembly。使用KMC工具(Kokot等,2017)对PacBio HiFi测序reads中的32碱基长度k-mer(k聚体)进行计数,并剔除出现次数超过10000次的k-mer。作为流程的一部分,我们基于KMC输出的k-mer频数直方图运行了GenomeScope工具(Ranallo-Benavidez等,2020),将其纳入方法部分以保证完整性。利用Smudgeplot工具(Ranallo-Benavidez等,2020)对样本的倍性水平进行分析。使用HiFiAdapterFilt工具(Sim等,2022)处理HiFi测序reads,以去除可能残留的PacBio接头序列。将过滤后的HiFi reads与Hi-C数据整合,通过Hifiasm工具(Cheng等,2021)进行组装,得到两套单倍型解析的基因组组装结果:假单倍型1(hap1)和假单倍型2(hap2)。利用meryl工具(Rhie等,2020)识别每套组装结果/假单倍型中的特异性k-mer,以此构建两组Hi-C reads:一组不含仅在hap1中出现的特异性k-mer,另一组不含仅在hap2中出现的特异性k-mer。将经过k-mer过滤的Hi-C reads与各自的支架化组装序列进行比对,使用BWA-MEM工具(Li,2013)并指定-5SPM参数。使用samtools工具(Li等,2009)按reads名称对比对结果进行排序,随后运行samtools fixmate命令以去除未比对reads和次级比对结果并添加mate评分,再运行samtools markdup命令去除重复reads。将生成的BAM文件用于两套组装结果的支架化步骤,使用YaHS工具(Zhou等,2023)并采用默认参数。使用FCS-GX工具(Astashyn等,2024)排查潜在污染序列,并移除污染序列。 lpNarStri2.hap1.decon.fasta.gz:硬羽茅(Nardus stricta)hap1假单倍型的基因组组装序列 lpNarStri2.hap2.decon.fasta.gz:硬羽茅(Nardus stricta)hap2假单倍型的基因组注释文件



