<b>The Hi-C scaffolding assembly (v1) and annotation of </b><b><i>P. citri</i></b>
收藏资源简介:
The Hi-C scaffolding assembly (v1) was constructed using the <i>P. citri</i> contig assembly (NCBI accession: GCF_014898815.1) and the Hi-C reads generated in this study.Low-quality raw reads (quality score <25, length shorter than 99 bp) and adaptors were removed using Fastp (version 0.23.2). The clean reads were then mapped to the contig assembly of P. citri (NCBI Accession number: GCF_014898815.1) using the Amira Genomics’ mapping pipeline (https://github.com/ArimaGenomics/mapping_pipeline). Briefly, the paired-end reads were mapped independently (as single-ends) using BWA-MEM (Version 0.7.17). For chimeric read, only the portion that maps in the 5’-orientation in relation to its read orientation was retained by using the script “filter_five_end.pl” that provided by Amira pipeline. Subsequently, these two filtered single-end Hi-C reads were then paired using the script “two_read_bam_combiner.pl”, which outputed a paired-end BAM file. Read group information was added to this paired-end BAM by using Picard module AddOrReplaceReadGroups. Any PCR duplicates present in the paired-end BAM file were removed by using Picard module MarkDuplicates. The valid paired-end pairs were then used for contig cluster, order and orient by yahs with different parameter settings. The interaction between contig pairs were converted into binary files by Juicer (v1.6). The Juicebox was employed to review assembly manually and generate the heat maps of contig interaction intensity and location.Liftoff (version 1.6.3) (https://github.com/agshumate/Liftoff) was used to accurately map annotation of previous published assembly to the new assembly generated in current study. Minimap2 was used to align the gene sequences from the previous published assembly to the new assembly generated in current study.
本研究基于柑橘生疫霉(*P. citri*)的重叠群(contig)组装序列(NCBI登录号:GCF_014898815.1)以及本研究产生的Hi-C测序读段,构建了Hi-C挂载组装版本v1。使用Fastp软件(版本0.23.2)去除低质量原始读段(质量分数<25、长度短于99 bp)与接头序列。随后,使用Amira Genomics公司的比对流程(https://github.com/ArimaGenomics/mapping_pipeline),将质控合格的洁净读段比对至柑橘生疫霉(*P. citri*)的重叠群组装序列(NCBI登录号:GCF_014898815.1)。简要而言,研究人员使用BWA-MEM软件(版本0.7.17)将双端读段以单端读段的形式独立进行比对。针对嵌合读段,使用Amira流程提供的"filter_five_end.pl"脚本,仅保留与读段方向一致的5’端比对区域对应的序列片段。随后,使用"two_read_bam_combiner.pl"脚本将这两个经过过滤的单端Hi-C读段重新组合为双端读段,最终生成双端BAM格式文件。使用Picard工具的AddOrReplaceReadGroups模块,为该双端BAM文件添加读段组信息。使用Picard工具的MarkDuplicates模块,去除该双端BAM文件中的PCR重复序列。将筛选得到的有效双端读段对用于yahs软件的不同参数设置下的重叠群聚类、排序与方向确定。使用Juicer软件(版本1.6)将重叠群对之间的相互作用数据转换为二进制文件。使用Juicebox软件对组装结果进行人工审核,并绘制重叠群相互作用强度与位置的热图。使用Liftoff软件(版本1.6.3,https://github.com/agshumate/Liftoff)将已发表参考组装的注释信息精准映射至本研究构建的新组装序列中。使用Minimap2软件将已发表参考组装中的基因序列比对至本研究构建的新组装序列中。




