Updated Spiny Mouse Transcriptome Assembly (Now Includes Embryo-Specific Transcripts)
收藏资源简介:
<strong>Summary</strong> Updated spiny mouse transcriptome. Embryo-specific contigs generated from BioProject PRJNA436818 were added to the Trinity_v2.3.2 spiny mouse <em>de novo </em>transcriptome assembly (https://doi.org/10.5281/zenodo.808870). <strong>Methods</strong> Embryos were collected from female spiny mice (n=12) in accordance with the Australian Code of Practice for the Care and Use of Animals for Scientific Purposes with approval from the Monash Medical Centre Animal Ethics Committee. Female dams were staged from delivery of their previous litter (spiny mice conceive their next litter approximately 12h postpartum) and culled at specific time-points for embryo retrieval at the required stage: 2-cell at 48h postpartum (n=4), 4-cell at 52h postpartum ('early' 4-cell; n=2) or at 68h postpartum ('late 4-cell'; n=2), and 8-cell at 72h postpartum (n=4). Embryos were snap frozen in cell lysis solution per the Nugen SoLo protocol (version M01406v3; available from NuGEN). After ligation of cDNA, qPCR was performed on all samples to determine the number of amplification cycles required to ensure that amplification was in the linear range. Based on these results, each sample was amplified using 24 cycles. Final libraries were quantitated by Qubit and size profile determined by the Agilent Bioanalyzer. All libraries were in the expected size range (~320-360 bp). Custom 'AnyDeplete' rRNA depletion probes were designed and produced by NuGEN Technologies, Inc (San Carlos, CA, USA) using rRNA sequences from our reference transcriptome (Mamrot et al., 2017; https://doi.org/10.5281/zenodo.808870). Prior to use, efficacy and off-target effects of the rRNA depletion probes were examined <em>in silico</em> by NuGEN. Samples were loaded using c-Bot (200pM per library pool) and run on 2 lanes of an Illumina HiSeq 3000 8-lane flow-cell. PhiX spike-in was not used directly due to incompatibility with the custom rRNA depletion probes, however it was incorporated into other lanes of the same HiSeq 3000 run. RNA-Seq data (100bp, paired-end reads) are available from the NCBI as Bioproject PRJNA436818. The quality of RNA-Seq reads was assessed using FastQC v0.11.6 (https://github.com/s-andrews/FastQC; 50f0c26), with MultiQC v1.4 (https://github.com/ewels/MultiQC; baefc2e) reports available from Github (https://github.com/jpmam1) (Ewels et al., 2016). Adapter sequences were trimmed from the reads using trim-galore v0.4.2 (https://github.com/FelixKrueger/TrimGalore; d6b586e), implementing cutadapt v1.12 (https://github.com/marcelm/cutadapt; 98f0e2f). Reads with a quality scores lower than 20 and read pairs in which either forward or reverse reads were trimmed to fewer than 35 nucleotides were discarded. Further trimming of poor quality reads was conducted using Trimmomatic v0.36 (http://www.usadellab.org/cms/index.php?page=trimmomatic) with settings "LEADING:3 TRAILING:3 SLIDINGWINDOW:4:20 AVGQUAL:25 MINLEN:35" (Bolger et al., 2014). Nucleotides with quality scores lower than 3 were trimmed from the 3’ and 5’ read ends. Reads with an average quality score lower than 25 or with a length of fewer than 35 nucleotides after trimming were removed. Error correction of reads was performed using Rcorrector v1.0.2 (https://github.com/mourisl/Rcorrector; 144602f) (Song et al., 2015). FastQC was used to assess the improvement in read quality after trimming adapter removal; MultiQC reports are available from Github (https://github.com/jpmam1). Error corrected reads were assembled using Trinity v2.4.0 (https://github.com/trinityrnaseq/trinityrnaseq; 1603d80) with settings "--max_memory 400G, --CPU 32 and --full_cleanup" (Haas et al., 2013). Assembly statistics were computed using the TrinityStats.pl from the Trinity package, and summary statistics are provided in Table S1. All reads were aligned to this transcriptome assembly using Bowtie2 v2.2.5 (https://github.com/BenLangmead/bowtie2; e718c6f) with settings: "--end-to-end, --score-min L,-0.1,-0.1, --no-mixed, --no-discordant, -k 100, -X 1000, --time, -p 24" (Langmead & Salzberg, 2012). Read-supported contigs were identified within the embryo-specific Trinity <em>de novo </em>transcriptome assembly using samtools "idxstats" v1.5 (contigs with >=1 reads aligning were retained) (https://github.com/samtools/samtools; f510fb1) (Li et al., 2009). The read-supported contigs from the embryo-specific assembly (n=54,660) were added to the reference spiny mouse transcriptome assembly previously described (Mamrot, J., Legaie, R., Ellery, S.J., Wilson, T., Seemann, T., Powell, D.R., Gardner, D.K., Walker, D.W., Temple-Smith, P., Papenfuss, A.T. and Dickinson, H., 2017. De novo transcriptome assembly for the spiny mouse (Acomys cahirinus). Scientific Reports, 7(1), p.8996). The updated transcriptome is comprised of 2,274,638 transcripts in total.



