遇见数据集

Draft assembly of Nardus stricta

收藏
Zenodo2026-06-02 更新2026-06-05 收录
官方服务:

资源简介:

We assembled the species using a pre-release of the EBP-Nor genome assembly pipeline (https://github.com/ebp-nor/GenomeAssembly). KMC (Kokot et al., 2017) was used to count k-mers of size 32 in the PacBio HiFi reads, excluding k-mers occurring more than 10,000 times. GenomeScope (Ranallo-Benavidez et al., 2020) was run as part of the pipeline on the k-mer histogram output from KMC and was included in the methods for completeness. Ploidy level was investigated using Smudgeplot (Ranallo-Benavidez et al., 2020). HiFiAdapterFilt (Sim et al., 2022) was applied on the HiFi reads to remove possible remnant PacBio adapter sequences. The filtered HiFi reads were assembled using Hifiasm (Cheng et al., 2021) with Hi-C integration resulting in a pair of haplotype-resolved assemblies, pseudo-haplotype one (hap1) and pseudo-haplotype two (hap2). Unique k-mers in each assembly/pseudo-haplotype were identified using meryl (Rhie et al., 2020) and used to create two sets of Hi-C reads, one without any k-mers occurring uniquely in hap1 and the other without k-mers occurring uniquely in hap2. K-mer filtered Hi-C reads were aligned to each scaffolded assembly using BWA-MEM (Li, 2013) with -5SPM options. The alignments were sorted based on name using samtools (Li et al., 2009) before applying samtools fixmate to remove unmapped reads and secondary alignments and to add mate score, and samtools markdup to remove duplicates. The resulting BAM files were used to scaffold the two assemblies using YaHS (Zhou et al., 2023) with default options. FCS-GX (Astashyn et al., 2024) was used to search for putative contamination and contaminated sequences were removed. lpNarStri2.hap1.decon.fasta.gz - genome assembly of Nardus stricta (hap1) lpNarStri2.hap2.decon.fasta.gz - genome annotation of Nardus stricta (hap2)

提供机构:
Zenodo
创建时间:
2026-06-02
二维码
社区交流群
二维码
科研交流群
商业服务