Psylliodes chrysocephala adult transcriptome annotation
收藏资源简介:
1.1 Library preparation <br> The total RNA from pre-aestivation (5-day-old), aestivation (30-day-old), and post-aestivation (55-day-old) female beetles were extracted using ZYMO Quick-RNA Tissue/Insect Kit (ZYMO Research, Irvine, CA, USA) and cleaned using TURBO DNA-free™ kit (Thermo Fisher Scientific, Langenselbold, Germany) according to the manufacturer’s instructions. We opted to sample only the females to eliminate sex-related variations. RNA quantity was determined using a Nanodrop ND-1000 UV/Vis spectrophotometer (Thermo Fisher Scientific). The integrity of the RNA samples was determined using the Agilent 2100 Bioanalyzer and an RNA 6000 Nano Kit (Agilent Technologies, Santa Clara, CA, USA). RIN values ≥ 7.0 were considered appropriate for mRNA library preparation. In total, 10 libraries (4, 3, and 3 libraries respectively per pre-aestivation, aestivation, and post-aestivation stages) were prepared using NEBNext® Poly(A) mRNA Magnetic Isolation Module kit (NEB E7490, New England Biolabs) according to the manufacturer’s instructions. The qualities of the libraries were checked via RNA fragment analysis conducted on the Agilent 2100 Bioanalyzer using the Agilent DNF-935 Reagent Kit (Agilent Technologies). The libraries were pooled based on their concentration, and an overall concentration of 3.4 ng/µL was obtained. The sequencing service was provided by BGI Genomics Tech Solutions Co. Ltd (Hong Kong) on a DNBSEQ-T7 platform. <br> 1.2 <em>De novo</em> assembly and functional annotation <br> Erroneous k-mers from paired read ends were removed using r-Corrector (v1.0.5) the with default options (Song & Florea, 2015), and the unfixable reads were discarded using the “FilterUncorrectabledPEfastq.py” function in Transcriptome Assembly Tools (Song & Florea, 2015). The adaptor sequences from the reads were removed, and the reads having a quality score above 30 were retained using TrimGalore! (v0.6.7). The cleaned reads (n = 3 per three adult phases) were <em>de novo</em> assembled using Trinity with default options. The completeness of the transcriptome was quantified using BUSCO (v5.4.2) via a comparison against the endopterygota dataset (BUSCO.v4 datasets). The transcriptome (including isoforms) was annotated using Trinotate (v3.2.2), which combines the outputs of NCBI BLAST+ (v2.13.0; nucleotide and predicted protein BLAST), TransDecoder (v5.5.0; coding region prediction), signal (v4.0; signal peptide prediction), TmHMM (v2.0; transmembrane domain prediction), and HMMER (v3.3.2; homology search) packages into an SQLite annotation database. The latest uniport_sprot (04/2022) and Pfam-A (11/2015) databases were downloaded using Trinotate, and the default E-value thresholds were used during the searches with BLAST+ and HMMER, respectively. The obtained annotation database was used to extract gene ontology (GO) terms associated with individual genes using the “extract_GO_assignments_from_Trinotate_xls.pl” whereas the signals and TmHMM outputs were manually extracted using Excel spreadsheets. The longest protein-coding regions in the super transcript data predicted by TransDecoder were subjected to Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway annotation via GhostKoala v2.2 (https://www.kegg.jp/ghostkoala/)
1.1 文库构建(Library preparation) 本研究采用ZYMO Quick-RNA组织/昆虫试剂盒(ZYMO Research,美国加利福尼亚州欧文市)提取夏眠前(5日龄)、夏眠期(30日龄)及夏眠后(55日龄)雌性甲虫的总RNA,并依照制造商说明书使用TURBO DNA-free™试剂盒(赛默飞世尔科技,德国朗根塞尔博尔德)完成RNA纯化。为消除性别相关差异,本研究仅选取雌性样本进行采样。使用Nanodrop ND-1000紫外-可见分光光度计(赛默飞世尔科技)测定RNA浓度,采用安捷伦2100生物分析仪配套RNA 6000 Nano试剂盒(安捷伦科技,美国加利福尼亚州圣克拉拉市)评估RNA完整性,RIN值≥7.0的样本被认定适用于mRNA文库构建。本研究共计构建10个文库:夏眠前、夏眠期、夏眠后阶段分别对应4、3、3个文库,均采用NEBNext® Poly(A) mRNA磁珠分离模块试剂盒(NEB E7490,新英格兰生物实验室),严格遵循制造商说明书完成构建。采用安捷伦DNF-935试剂试剂盒(安捷伦科技)在安捷伦2100生物分析仪上开展RNA片段分析以验证文库质量。依照文库浓度进行混样,最终获得总浓度为3.4 ng/µL的混合文库。测序服务由华大基因科技解决方案有限公司(香港)提供,测序平台为DNBSEQ-T7。 1.2 从头(de novo)组装与功能注释 使用r-Corrector(v1.0.5)以默认参数去除配对末端读段中的错误k-mer(Song & Florea, 2015),并通过转录组组装工具中的"FilterUncorrectabledPEfastq.py"函数丢弃无法校正的读段。使用TrimGalore!(v0.6.7)去除读段中的接头序列,并保留质量分数≥30的读段。针对三个成虫阶段各3个清洁后读段样本,采用Trinity以默认参数进行从头组装。采用BUSCO(v5.4.2)比对内翅类数据集(BUSCO.v4数据集)以评估转录组的完整性。使用Trinotate(v3.2.2)对转录组(包括可变剪接异构体)进行注释,该工具将NCBI BLAST+(v2.13.0;核苷酸及预测蛋白BLAST)、TransDecoder(v5.5.0;编码区预测)、信号肽预测工具(SignalP v4.0)、TmHMM(v2.0;跨膜结构域预测)及HMMER(v3.3.2;同源性搜索)的输出结果整合至SQLite注释数据库中。使用Trinotate下载最新的UniProt_Swissprot(2022年4月)及Pfam-A(2015年11月)数据库,BLAST+与HMMER搜索分别采用默认E值阈值。使用"extract_GO_assignments_from_Trinotate_xls.pl"脚本从所得注释数据库中提取单个基因对应的基因本体(Gene Ontology, GO)术语,而信号肽预测工具及TmHMM的输出结果则通过Excel电子表格手动提取。将TransDecoder预测的超级转录本数据集中最长的蛋白编码序列,通过GhostKoala v2.2(https://www.kegg.jp/ghostkoala/)进行京都基因与基因组百科全书(Kyoto Encyclopedia of Genes and Genomes, KEGG)通路注释。



