Curated protein database for Gossypium hirsutum proteomics (seed and stress-related)
收藏资源简介:
This dataset contains a curated and non-redundant protein FASTA database for Gossypium hirsutum, optimized for shotgun proteomics analysis. The database was constructed by:- Downloading the complete UniProtKB protein dataset for Gossypium hirsutum- Filtering sequences shorter than 50 amino acids- Removing exact duplicates- Reducing redundancy using CD-HIT at 95% identity- Enriching the dataset with proteins related to seed development and abiotic stress (stress response, oxidative stress, salt stress, osmotic stress)- Removing redundancy in enriched subsets (CD-HIT at 90%)- Adding common proteomics contaminants (keratins, trypsin) Final database size:49,878 protein sequences This database is intended for use in mass spectrometry-based proteomics workflows (e.g., PLGS, Proteome Discoverer).



