Annotated protein (amino acid) sequences of the common woodlouse Porcellio scaber (Crustacea, Isopoda)
收藏资源简介:
This dataset contains annotated protein (amino acid) sequences of the common woodlice Porcellio scaber (Crustacea: Isopoda), an important model species in ecotoxicology, immunology, and physiological research. In this context, the dataset serves as a cornerstone for future proteome-based research on P. scaber. The assembled transcriptome of P. scaber was annotated in several steps. First, it was aligned to the SwissProt database using BLASTP from ncbiblastplus v2.14.1 (Camacho et al., 2009). To use only arthropod data, the -taxidlist option was used with the tax ID list from NCBI for arthropods. To limit the amount of output, the options -max_target_seqs 3 -evalue 1e-6 were used. Next, the results from the first search were filtered for sequences that did not have sufficiently good hits using a custom script. This script filters the results of the BLAST search and searches the fasta file for sequences without adequate hits. Using seqtk v1.4 (https://github.com/lh3/seqtk) subseq, the original fasta file was filtered for sequences not found in the SwissProt BLAST search. In the next step, InterProScan (Jones et al., 2014) was used on the sequences that were not identified by the previous search. In a similar fashion, a custom bash script was used to filter the input fasta file so that only the sequences not identified by this and the previous step remained. These sequences were used as the input for another BLAST search against the nr database, with the other settings the same as before. The final fasta file was created using a custom R script. The headers in the final fasta file contain first the sample name and the sequence number in the file, then the method used for annotating the sequence, and finally the description provided by the method. Unidentified sequences are marked with NA.



