Chromosome-scale genome assembly of the African spiny mouse (Acomys cahirinus)
收藏资源简介:
Genomic DNA was extracted from blood from a single male A. cahirinus animal using a Monarch HMW DNA Extraction Kit for Cells & Blood (T3050, New England Biolabs, Ipswich MA) following the manufacturer’s recommended protocol. DNA was quantified prior to library construction using the Qubit DNA HS Assay (ThermoFischer, Waltham MA) and DNA fragment lengths were assessed using the Agilent Femto Pulse System (Santa Clara, CA). Libraries were prepared for sequencing using the Oxford Nanopore ligation kit (SQK-LSK110) following the manufacturers’ instructions, except that DNA repair and A-tailing was performed for 30 min and the ligation was allowed to continue for 1 hr. Prepared libraries were quantified using a Qubit fluorometer and 30 fmol of the library was loaded onto a Nanopore version R.9.4.1 flow cell and loaded on a PromethION running MinKNOW version (21.05.20). To increase output, the flow cell was washed after approximately 24 hr of sequencing then an additional 12 fmol of library was added to the flow cell and run for an additional 48 hr. Basecalling was performed using Guppy 5.0.12 (Oxford Nanopore) using the superior model (dna_r9.4.1_450bps_sup_prom.cfg). FASTQ files for assembly were extracted from unaligned bam files using samtools (Li et al. 2009) then Flye version 2.9 for assembly using the --nano-hq flag (Kolmogorov et al. 2019). Haplotigs and overlaps in the assembly were purged using purge_dups (https://github.com/dfguan/purge_dups). The assembly was then polished using Medaka version 1.4.2 (https://github.com/nanoporetech/medaka) followed by a second polishing step with pilon version 1.24 (Walker et al. 2014). Assembly statistics at each step were generated using Quast (Gurevich et al. 2013) and BUSCO (Simão et al. 2015) (Table S2). The primary contigs assembled from the Nanopore data were anchored to chromosomes using 505,210,505 read pairs of a Hi-C library isolated from another A. cahirinus individual of unknown sex downloaded from the NCBI Short Read Archive (SRX13258644) (Wang et al. 2022). After aligning the Hi-C reads with the ArimaHi-C Mapping Pipeline (https://github.com/ArimaGenomics/mapping_pipeline), YaHS v1.0 (Zhou et al. 2023) was used with default error correction for scaffolding, and Juicebox v1.11.08 (Dudchenko et al. 2018) was used to generate a Hi-C contact map. Progressive Cactus was used (Armstrong et al. 2020) to perform a whole-genome alignment of the A. cahirinus draft assembly to the Mus musculus GRCm39 reference genome (RefSeq GCF_000001635.27_GRCm39). Comparative annotation of the draft genomes was then performed using the Comparative Annotation Toolkit (CAT) (Fiddes et al. 2018). Briefly, the M. musculus RefSeq annotation GFF was parsed and validated with the “parse_ncbi_gff3” and “validate_gff3” programs (respectively) from CAT. The M. musculus reference transcript cDNA sequences were downloaded and mapped to the M. musculus draft genome with minimap2 (Li 2018) and provided to CAT as long-read RNA-seq reads in the “[ISO_SEQ_BAM]” field of the configuration file. For A. cahirinus, bulk RNA-seq data obtained from multiple pooled organs were downloaded from NCBI SRA BioProject PRJNA342864 (Bellofiore et al. 2017) and mapped to the draft assembly with STAR (Dobin et al. 2013) then provided to CAT in the “[BAMS]” field. CpG islands were identified using the cpg_lh utility from the UCSC suite of tools (Kent et al. 2002).
本研究从单只雄性埃及刺毛鼠(A. cahirinus)的血液中提取基因组DNA:采用细胞与血液高分子量DNA提取试剂盒(Monarch HMW DNA Extraction Kit for Cells & Blood, 货号T3050, 纽英伦生物实验室, 美国马萨诸塞州伊普斯维奇),按照制造商推荐的实验流程完成提取。文库构建前,使用Qubit DNA高灵敏度检测试剂盒(Qubit DNA HS Assay, 赛默飞世尔科技, 美国马萨诸塞州沃尔瑟姆)对DNA进行定量,并通过安捷伦Femto脉冲系统(Agilent Femto Pulse System, 美国加利福尼亚州圣克拉拉)评估DNA片段长度。采用牛津纳米孔连接试剂盒(Oxford Nanopore ligation kit, SQK-LSK110),按照制造商的操作指南构建测序文库,相较于标准流程,本研究将DNA修复与A尾加尾反应时长调整为30分钟,连接反应时长延长至1小时。构建完成的文库通过Qubit荧光定量仪进行定量,取30飞摩尔(fmol)的文库加载至纳米孔R.9.4.1版本流动槽,搭载MinKNOW 21.05.20版本软件的PromethION测序仪启动测序。为提升测序产出,测序约24小时后对流动槽进行清洗,随后补加12飞摩尔文库至流动槽,继续测序48小时。碱基识别(basecalling)步骤采用Guppy 5.0.12软件(牛津纳米孔公司),使用高精度模型dna_r9.4.1_450bps_sup_prom.cfg完成。使用samtools从未比对的BAM文件中提取用于基因组组装的FASTQ文件,随后采用Flye 2.9版本软件,以--nano-hq参数进行基因组组装。使用purge_dups工具(https://github.com/dfguan/purge_dups)去除组装结果中的单倍型片段与重叠序列。组装结果先通过Medaka 1.4.2版本软件(https://github.com/nanoporetech/medaka)进行抛光校正,再使用Pilon 1.24版本软件(Walker et al. 2014)完成第二轮抛光校正。使用Quast(Gurevich et al. 2013)与BUSCO(Simão et al. 2015)软件生成各步骤的组装统计数据,结果详见附表S2。从NCBI短读长档案库(NCBI Short Read Archive, SRX13258644)下载另一只性别未知的埃及刺毛鼠个体的Hi-C文库(Hi-C library)数据(共505,210,505条读对),将其用于将纳米孔测序组装得到的主要重叠群(contigs)锚定至染色体(Wang et al. 2022)。将Hi-C读段与ArimaHi-C比对流程(https://github.com/ArimaGenomics/mapping_pipeline)完成比对后,采用YaHS v1.0软件(Zhou et al. 2023)以默认参数进行错误校正并构建基因组支架,再使用Juicebox v1.11.08软件(Dudchenko et al. 2018)生成Hi-C相互作用图谱。采用Progressive Cactus工具(Armstrong et al. 2020)将埃及刺毛鼠的草稿组装基因组与小家鼠(Mus musculus)GRCm39参考基因组(RefSeq GCF_000001635.27_GRCm39)进行全基因组比对。随后使用比较注释工具包(Comparative Annotation Toolkit, CAT)(Fiddes et al. 2018)对草稿基因组进行比较注释:简要而言,使用CAT工具包中的"parse_ncbi_gff3"与"validate_gff3"程序分别对小家鼠RefSeq注释的GFF文件进行解析与验证;下载小家鼠参考转录本cDNA序列,使用minimap2(Li 2018)将其比对至小家鼠草稿基因组,并将其作为长读长RNA-seq读段,以配置文件的"[ISO_SEQ_BAM]"字段形式提供给CAT工具包。针对埃及刺毛鼠,从NCBI SRA生物项目PRJNA342864(Bellofiore et al. 2017)下载多器官混合的批量RNA-seq数据,使用STAR软件(Dobin et al. 2013)将其比对至埃及刺毛鼠的草稿组装基因组,并以"[BAMS]"字段形式提供给CAT工具包。最后使用UCSC工具集中的cpg_lh工具(Kent et al. 2002)识别基因组中的CpG岛。



