Dientamoeba fragilis transcriptome assembly
收藏资源简介:
Re-assembly and bacterial contamination removal from the Dientamoeba fragilis transcriptome first published in: Barratt JLN, Cao M, Stark DJ, Ellis JT. The Transcriptome Sequence of Dientamoeba fragilis Offers New Biological Insights on its Metabolism, Kinome, Degradome and Potential Mechanisms of Pathogenicity. Protist. 2015;166:389–408. https://doi.org/10.1016/j.protis.2015.06.002 Methods: RNA-seq reads were downloaded from the SRA under accession SRR2039085 and assembled using Trinity version 2.15.1. Protein-coding regions were identified using TransDecoder version 5.7.1, resulting in 38,220 protein-coding regions sequences (https://github.com/TransDecoder/TransDecoder). To identify and remove sequences that may have originated from bacteria in the xenic culture with Dientamoeba fragilis, we scanned the predicted proteins with the alien inde. All proteins (translated CDS regions) were aligned against the clustered NR database using diamond version 2.1.9.163, resulting in 28,586 proteins with a significant hit. The alien index was calculated as AI=bbsO/bbsS - bbsI/bbsS, where bbsO is the bit score of the best hit to a species outside of the group lineage, bbsI is the bit score of the best hit to a species within the group lineage (skipping self hits), and bbsS is the bit score of the query aligned to itself. The recipient lineage was designated as eukaryotes. Genes with an alien index below 0.1 were considered likely eukaryotic in origin and retained. While this likely excludes genes that are truly present in D. fragilis genome, this conservative approach prevents bacterial sequences from entering the EukDetect2 database. Protein-coding regions were clustered at 99% identity to remove duplicates using cd-hit version 4.8.1, resulting in 11,404 protein-coding regions.



