Comprehensive Eukaryotic Proteomes (200+ aa) from UniProtKB Swiss-Prot and TrEMBL
收藏资源简介:
Description This dataset contains a comprehensive collection of eukaryotic protein sequences retrieved from the UniProt Knowledgebase (UniProtKB), combining both reviewed (Swiss-Prot) and unreviewed (TrEMBL) entries. The dataset was generated using a specific filtering pipeline designed to capture functional eukaryotic proteomes while excluding short fragments and small peptides, making it ideal for large-scale comparative genomics, phylogenetics, hidden Markov model (HMM) profile searches, and deep learning protein language model training. Dataset Specifications & Query Parameters The data was fetched from the official UniProtKB database using the following precise search query: (taxonomy_id:2759) AND (length:[200 TO *]) Taxonomic Scope: Eukaryota (Taxonomy ID: 2759) — includes all major eukaryotic lineages (Metazoa, Viridiplantae, Fungi, and various protist groups). Sequence Length Filter: Only proteins containing 200 or more amino acids (aa) are included ($\ge 200$ aa). This threshold effectively filters out truncated sequences, low-quality fragments, and short structural peptides. Database Coverage: Combined Swiss-Prot + TrEMBL (complete UniProtKB snapshot). File Format & Contents Format: Standard FASTA format (.fasta), compressed using gzip (.gz). Identifiers: Standard UniProt headers containing Accession Number, Entry Name, Protein Name, Organism Name (OS=), Organism Identifier (OX=), Gene Name (GN=), Protein Existence (PE=), and Sequence Version (SV=). Potential Applications This dataset is optimized for: High-throughput sequence similarity searches (e.g., using BLAST, DIAMOND, or HMMER). Constructing customized sequence databases for specific eukaryotic phila. Training machine learning models requiring a representative and comprehensive set of large eukaryotic proteins How to Reassemble the Dataset Due to Zenodo web-upload size limitations, this 19 GB archive has been split into smaller 300 MB parts. To reconstruct the original .gz file on macOS or Linux, download all parts into the same directory, open your terminal, and run: cat uniprot-eukaryota-proteomes-200-plus.fasta.gz.part_* > uniprot-eukaryota-proteomes-200-plus.fasta.gz To ensure the file was reassembled correctly and without corruption, you can verify its SHA-256 checksum. Run the following command in your terminal: shasum -a 256 uniprot-eukaryota-proteomes-200-plus.fasta.gz Expected SHA-256: 50081d0fe6d3ae92b3b5f37a780ea7cd848c3c06f635491cc4a7c7b11376a10b All protein sequences were retrieved from the UniProt Knowledgebase (UniProt Consortium). Distributed under the terms of the Creative Commons Attribution (CC BY 4.0) License.



