<b>The Hi-C scaffolding assembly (v1) and annotation of </b><b><i>P. citri</i></b>
收藏资源简介:
The Hi-C scaffolding assembly (v1) was constructed using the <i>P. citri</i> contig assembly (NCBI accession: GCF_014898815.1) and the Hi-C reads generated in this study.Low-quality raw reads (quality score <25, length shorter than 99 bp) and adaptors were removed using Fastp (version 0.23.2). The clean reads were then mapped to the contig assembly of P. citri (NCBI Accession number: GCF_014898815.1) using the Amira Genomics’ mapping pipeline (https://github.com/ArimaGenomics/mapping_pipeline). Briefly, the paired-end reads were mapped independently (as single-ends) using BWA-MEM (Version 0.7.17). For chimeric read, only the portion that maps in the 5’-orientation in relation to its read orientation was retained by using the script “filter_five_end.pl” that provided by Amira pipeline. Subsequently, these two filtered single-end Hi-C reads were then paired using the script “two_read_bam_combiner.pl”, which outputed a paired-end BAM file. Read group information was added to this paired-end BAM by using Picard module AddOrReplaceReadGroups. Any PCR duplicates present in the paired-end BAM file were removed by using Picard module MarkDuplicates. The valid paired-end pairs were then used for contig cluster, order and orient by yahs with different parameter settings. The interaction between contig pairs were converted into binary files by Juicer (v1.6). The Juicebox was employed to review assembly manually and generate the heat maps of contig interaction intensity and location.Liftoff (version 1.6.3) (https://github.com/agshumate/Liftoff) was used to accurately map annotation of previous published assembly to the new assembly generated in current study. Minimap2 was used to align the gene sequences from the previous published assembly to the new assembly generated in current study.



