遇见数据集

Genome and developmental transcriptome of the forensically important blowfly, <em>Phormia Regina</em>

收藏
NIAID Data Ecosystem2026-05-10 收录
官方服务:

资源简介:

Forensic entomologists often rely on the morphological traits of insects, such as length and weight, to estimate the time of death. Blowflies like Phormia regina are particularly significant in North American forensic investigations. However, age estimation of older P. regina maggots becomes challenging due to limited morphological changes during the lengthy L3 larval phase of development. To address this gap, we used transcriptomic profiling of blowfly maggots to generate molecular markers that specify their age. We first characterized maggot weight, behavior, and mRNA in 10-hour increments during development. At 27.5°C, the weight of the maggots increased from when first recorded at 70 through 100 hours and then remained stable from 110 hours to pupation. The behavioral transition between the feeding and the post-feeding wandering stage usually took place between 100 and 120 hours. Second, we built a chromosomal-scale Phormia regina genome annotated with long mRNA reads to provide a reliable database to uncover transcriptomic signatures during larval development. We applied differential gene expression analysis (DEGs), weighted gene co-expression network analysis (WGCNA), and the generalized linear model (GLM) to identify nine candidate genes that all three statistical analyses indicated are useful for delineating the age of otherwise indeterminate L3 maggots. In turn, these genes could be used to design a quantitative PCR protocol for more accurate estimates of the time of death. Methods Identifying the transition between feeding and wandering behavior We maintained a laboratory colony of a highly inbred line of Phormia regina. The source was from Dr. Amanda Roe at University of Nebraska at Lincoln [1]. Developmental rate data were available [1] . All rearing containers were maintained at 27.5+ 0.5 °C. under 16:8 light:dark cycle in one SMY04-1 DigiTherm CirKinetics Incubator (TriTech Research, Inc., Los Angeles, California, USA). The incubator was equipped with uniform lighting, additional fans, and a port for thermometer access. Each equal-age cohort was reared within a plastic 72 x72x100 mm insect breeding box (Gyeonggi-do, Korea). Each box contained 0.5cm sawdust and a 30g of fresh chicken liver in a suspended 4 oz paper cup. Eggs were obtained on a paper towel soaked with chicken liver blood placed in a cage of P. regina adults for 30 minutes before removal. The removal time was considered age zero for those insects. About 500-1000 newly deposited eggs were placed on wet paper in a covered petri dish. The eggs were kept at high humidity over an open water container at 27.5+ 0.5 °C. After 24 hours, fifteen newly hatched maggots were transferred into each aliquot of fresh liver inside the rearing box (See Fig 1A). All the boxes were reared at the same 27.5+ 0.5 °C incubator. For each of four trials, 28 boxes with maggots were set up. During each sampling time, four rearing boxes were randomly removed from the incubators. We distinguished a maggot as feeding if it was still on the liver and wandering if it was in the sawdust. Sampling involved removal of an entire cohort at a preselected age. Sampled maggot was individually placed in a 1.5mL microcentrifuge tube with 1mL of Thermofisher RNAlater (Invitrogen, MA, USA) at 4 °C. Each tube was labeled with numbers for RNA analysis. After 24 hours, the RNAlater was removed from the tube and the insect was stored at -80 °C. Build a new P. regina genome assembly To create a high-quality reference assembly for P. regina, a single male adult was sequenced by PacBio Sequel II and we performed de novo assembly using HiFiasm v0.16.1 with purge mode to generate a male P. regina haplotype and diplotype assembly [2] . Duplicate contigs were further eliminated by purge_dups v1.2.6 [3]. The contaminated contigs were identified by NCBI BLAST against a publicly available nucleotide database and filtered by Blobtool v2.6.4[4] . The repeat element boundaries and repeat database was de novo assembly by the clean contigs using RepeatModeler v2.0.2 [5–13]. A single male adult fly was sequenced by Illumina NovaSeq 6000 to create Hi-C sequencing libraries. Hi-C sequencing libraries were first mapped to the clean contigs by Arima mapping and the .bam file was used to perform chromosomal integrated assembly by yahs v1.1 [14] . A Hi-C heat map was generated by juicebox v2.20.00 [15]. The curation and chromosomal boundaries were manually edited following the Genome Assembly Cookbook [16]. The chromosomal genome was annotated using the Funannotate v1.8.13 annotation pipeline, with Iso-Seq libraries providing support for evidence-based gene prediction during the annotation process. (See Isoseq mRNA analysis of P. regina ) [17–22]. RNA extraction For each age cohort, one labeled maggot was selected based on the average weight of cohorts. The specimen was homogenized using pestle with 200 ul of ambion TRIzol reagent (MA, USA). RNA was extracted from using a QIAGEN RNeasy Plus extraction kit (MA, USA). The concentration of the RNA samples was assessed with an Invitrogen Qubit Fluorometer using a Qubit RNA HS Assay kit (MA, USA). The integrity of the extracted RNA was evaluated with a Bioanalyzer using an Agilent RNA 6000 Pico chip. The process was repeated three times for 3 replicates for each age cohorts (Hong Kong, China). A total of 33 RNA samples from maggots of known ages were sent to BGI (CA, USA) for Illumina RNA sequencing, and 2 RNA samples from adult flies were sent to Genewiz (NJ, USA). Isoseq mRNA analysis of P. regina A single virgin adult male, a single juvenile virgin adult female, a 110 hour feeding third instar maggot, and a 110 hour wandering maggot were submitted to PacBio (CA, USA) for Iso-Seq on the PacBio Sequel II platform. Consensus sequences were generated from the SMRTBell libraries and collapsed into isoforms by Isoseq v3.0 [23]. The long reads libraries were used as evidence for genome annotation. Gene expression analysis The quality of RNA seq libraries were initially assessed by FastQC v0.11.9 [24]. The adapter sequences from the short reads were trimmed by Trimmomatic v0.39 [25]. The clean short read libraries were mapped to the newly annotated genome using Salmon v1.8.0 for quantification [26]. Count matrices were obtained by tximport [27]. Differential expression analysis was performed by DeSeq2 1.34.0 Bioconductor [28]. The gene names were manually curated based on the results from the BLASTn against orthologous UniProt v2023_02. Pairwise comparison was performed among all treatments to identify differentially expressed transcripts (> 2-fold change, a < 0.01 FDR) using DEBrowser v3.18[29]. Gene co-expression analysis was produced by DESeq2 v1.34.0 followed by tidyverse v2.0.0, magrittr v2.0.3, and WGCNA v1.72.1 analysis on R v2022.12.0+353 [30–31]. Gene ontology enrichment analysis was performed by using ShinyGOV0.76 under FDR-corrected cut off a < 0.05 [32]. Finding candidate genes for predicting larval age. Raw reads from the development of aging maggot were normalized to TPM values by DeSeq2 1.34.0 [28]. Linear regression model was applied to define the housekeeping genes (s <-2, all values greater than 0), upregulated genes ( b1 > 1 and R2: 0.9~1) and downregulated genes ( b1 < -1 and R2: 0.9~1) from the mean TPM values across each age cohort. The correlation plots were created using ggplot2 [33] The Venn diagram list was created by the VennDiagram v1.7.3 package from RStudio 0.15.0 and drawn by Procreate v5.3.7 from Savage Interactive Pty Ltd [34-35] . Linear regression figures were created by Prism10 v10.2.2 from GraphPad software. 1. Roe AL. Development Modeling of Lucilia Sericata and Phormia Regina (Diptera : Calliphoridae). University of Nebraska-Lincoln; 2014. 2. Cheng H, Concepcion GT, Feng X, Zhang H, Li H. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nat Methods. 2021;18: 170–175. 3. Guan D, McCarthy SA, Wood J, Howe K, Wang Y, Durbin R. Identifying and removing haplotypic duplication in primary genome assemblies. Bioinformatics. 2020;36: 2896–2898. 4. Challis R, Richards E, Rajan J, Cochrane G, Blaxter M. BlobToolKit – Interactive Quality Assessment of Genome Assemblies. G3 Genes|Genomes|Genetics. 2020;10: 1361–1374. 5. Flynn JM, Hubley R, Goubert C, Rosen J, Clark AG, Feschotte C, et al. RepeatModeler2 for automated genomic discovery of transposable element families. Proc Natl Acad Sci U S A. 2020;117: 9451–9457. 6. Bao Z, Eddy SR. Automated de novo identification of repeat sequence families in sequenced genomes. Genome Res. 2002;12: 1269–1276. 7. Price AL, Jones NC, Pevzner PA. De novo identification of repeat families in large genomes. Bioinformatics. 2005;21 Suppl 1: i351–8. 8. Benson G. Tandem repeats finder: a program to analyze DNA sequences. Nucleic Acids Res. 1999;27: 573–580. 9. Ellinghaus D, Kurtz S, Willhoeft U. LTRharvest, an efficient and flexible software for de novo detection of LTR retrotransposons. BMC Bioinformatics. 2008;9: 18. 10. Ou S, Jiang N. LTR_retriever: A highly accurate and sensitive program for identification of long terminal repeat retrotransposons. Plant Physiol. 2018;176: 1410–1422. 11. Katoh K, Misawa K, Kuma K-I, Miyata T. MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform. Nucleic Acids Res. 2002;30: 3059–3066. 12. Li W, Godzik A. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. Bioinformatics. 2006;22: 1658–1659. 13. Wheeler TJ. Large-Scale Neighbor-Joining with NINJA. Algorithms in Bioinformatics. Algorithms in Bioinformatics. 2009; 375–389. 14. Zhou C, McCarthy SA, Durbin R. YaHS: yet another Hi-C scaffolding tool. Bioinformatics. Edited by C. Alkan; 2023. 15. Durand NC, Robinson JT, Shamim MS, Machol I, Mesirov JP, Lander ES, et al. Juicebox Provides a Visualization System for Hi-C Contact Maps with Unlimited Zoom. Cell Syst. 2016;3: 99–101. 16. Neva C. Durand, Muhammad S. Shamim, Ido Machol, Suhas S. P. Rao, Miriam H. Huntley, Eric S. Lander, Olga Dudchenko, Sanjit S. Batra, Arina D. Omer, Sarah K. Nyquist, Marie Hoeger, Neva C. Durand, Muhammad S. Shamim, and Erez Lieberman Aiden. Genome Assembly Cookbook. THE CENTER FOR GENOME ARCHITECTURE. Baylor College of Medicine & Rice University. 17. Palmer JM. Funannotate: pipeline for genome annotation. 2016. 18. Stanke M, Morgenstern B. AUGUSTUS: a web server for gene prediction in eukaryotes that allows user-defined constraints. Nucleic Acids Res. 2005;33: W465–7. 19. Bromberg Y, Rost B. SNAP: predict effect of non-synonymous polymorphisms on function. Nucleic Acids Res. 2007;35: 3823–3835. 20. Haas BJ, Salzberg SL, Zhu W, Pertea M, Allen JE, Orvis J, et al. Automated eukaryotic gene structure annotation using EVidenceModeler and the Program to Assemble Spliced Alignments. Genome Biol. 2008;9: R7. 21. Grabherr MG, Haas BJ, Yassour M, Levin JZ, Thompson DA, Amit I, et al. Full-length transcriptome assembly from RNA-Seq data without a reference genome. Nat Biotechnol. 2011;29: 644–652. 22. Haas BJ, Delcher AL, Mount SM, Wortman JR, Smith RK Jr, Hannick LI, et al. Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies. Nucleic Acids Res. 2003;31: 5654–5666. 23. Zhou G, Chen X, Pang J, Srinives P. Domestication of Agronomic Traits in Legume Crops. Frontiers Media SA; 2021. 24. Andrews S. FastQC: a quality control tool for high throughput sequence data. 2010. 25. Bolger AM, Lohse M, Usadel B. Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics. 2014;30: 2114–2120. 26. Patro R, Duggal G, Love MI, Irizarry RA, Kingsford C. Salmon provides fast and bias-aware quantification of transcript expression. Nat Methods. 2017;14: 417–419. 27. Soneson C, Love MI, Robinson MD. Differential analyses for RNA-seq: transcript-level estimates improve gene-level inferences. F1000Res. 2015;4: 1521. 28. Love MI, Huber W, Anders S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 2014;15: 550. 29. Kucukural A, Yukselen O, Ozata DM, Moore MJ, Garber M. DEBrowser: interactive differential expression analysis and visualization tool for count data. BMC Genomics. 2019;20: 6. 30. Langfelder P, Horvath S. WGCNA: an R package for weighted correlation network analysis. BMC Bioinformatics. 2008;9: 559. 31. Bache SM, Wickham H. magrittr: a forward-pipe operator for R. R package version. 32. Ge SX, Jung D, Yao R. ShinyGO: a graphical gene-set enrichment tool for animals and plants. Bioinformatics. 2020;36: 2628–2629. 33. Wickham H. Getting Started with ggplot2. In: Wickham H, editor. ggplot2: Elegant Graphics for Data Analysis. Cham: Springer International Publishing; 2016. pp. 11–31. 34. Wickham H. Getting started with ggplot2. Use R! Cham: Springer International Publishing; 2016. pp. 11–31. 35. R Core Team, R. R: A language and environment for statistical computing. 2013.

创建时间:
2025-10-27
二维码
社区交流群
二维码
科研交流群
商业服务