LukProt - an animal evolution-centric eukaryotic protein database
收藏资源简介:
LukProt is the EukProt database with additional species added, mostly the undersampled animal and holozoan taxa. The database is composed of sequences translated from annotated genomes, transcriptomes or ESTs. The main purpose of the database is to be a resource to look for whether a given protein or domain is present in large clades and to reconstruct its pedigree. The current version of the database (v1.4.1) is based on EukProt v2. The home of this database will be Zenodo. The EukProt species identifiers (EPXXXXX, where XXXXX is numbering, starting from 00001) and the dataset addition philosophy was conserved: each dataset gets its own identifier. Those that are novel in LukProt are denoted as LPXXXXX. Additionally, each sequence was assigned an ID in the following format: <pre><code>EPXXXXX_Species_epithet_(strain).#.</code></pre> where # is a single GI-like number assigned to each sequence within a taxon. All the IDs are compatible with BLAST v5 parse_seqids option and the database can be readily deployed, for example on a server running SequenceServer. Within each of the source fasta files, the source sequence identifier was kept after a blank space, so that it can still be retrieved if needed. Comparison of EukProt v2 and LukProt v1.4.1 in their main areas of difference: Taxogroup EukProt v2 LukProt v1.4.1 Unicellular Holozoa 31 39 Porifera 4 30 Ctenophora 2 35 Placozoa 2 3 Cnidaria 3 65 Bilateria 53 94 Included with the database are: a FASTA file with the sequences - 7-zipped a preformatted BLAST database (v5, masked with segmasker) - 7-zipped a file for taxon-based tree coloring (using figtree-recolor) a spreadsheet with information about each dataset (in an open .ods format, most compatible with LibreOffice) a pdf with the phylogeny and coloring scheme, as well as ID lists for taxogroup filtering a README file Words of caution: Many datasets, especially those transcriptome-based, may contain contamination from different species. In addition, the translation algorithms often introduce errors (e.g. the transcript is not full). For this reason, to get accurate sequences from each organism, the users are directed to source data. The taxonomy string for each taxon is often not proper. Some NCBI taxids are missing. The phylogeny/coloring scheme needs to be drawn more nicely. One dataset (LP00008), present in some metadata, is unpublished and has been held back. While the database contains metadata that present a particular phylogeny of animals, holozoans and other eukaryotes, no particular claims or hypotheses are made by the author(s). The names "Placnidia" and "Bilplacnidia" were invented by the author for convenience and are not in any way accepted in the community. EukProt has already been updated to a newer version (v3). This database will be updated to new EukProt versions after some time. Acknowledgements: Andrew E. Allen Lab for creating the original PhyloDB. Daniel Richter <em>et al.</em> for creating EukProt and keeping it updated. Members of the Multicellgenome Lab, especially Michelle Leger (for donating her database), for the bioinformatics support and for doing great science. All the authors of the original datasets. National Science Centre of Poland for funding of the project 2020/36/C/NZ8/00081, "The role of glycosylation in the emergence of animal multicellularity", which enabled the creation of this database.



