遇见数据集

Curated protein database for Gossypium hirsutum proteomics (seed and stress-related)

收藏
Zenodo2026-04-30 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains a curated and non-redundant protein FASTA database for Gossypium hirsutum, optimized for shotgun proteomics analysis. The database was constructed by:- Downloading the complete UniProtKB protein dataset for Gossypium hirsutum- Filtering sequences shorter than 50 amino acids- Removing exact duplicates- Reducing redundancy using CD-HIT at 95% identity- Enriching the dataset with proteins related to seed development and abiotic stress (stress response, oxidative stress, salt stress, osmotic stress)- Removing redundancy in enriched subsets (CD-HIT at 90%)- Adding common proteomics contaminants (keratins, trypsin) Final database size:49,878 protein sequences This database is intended for use in mass spectrometry-based proteomics workflows (e.g., PLGS, Proteome Discoverer).

本数据集包含一套经过精心甄选且无冗余的陆地棉(Gossypium hirsutum)蛋白质FASTA数据库,专为鸟枪法蛋白质组学分析优化。 该数据库的构建流程如下: 1. 下载陆地棉的完整UniProtKB蛋白质数据集; 2. 过滤掉长度小于50个氨基酸的序列; 3. 移除完全重复的序列; 4. 使用CD-HIT以95%的序列同一性阈值降低数据库冗余度; 5. 向数据集中补充与种子发育及非生物胁迫(涵盖胁迫响应、氧化胁迫、盐胁迫、渗透胁迫)相关的蛋白质; 6. 对富集后的子集再次使用CD-HIT,以90%的序列同一性阈值去除冗余; 7. 添加常见的蛋白质组学污染物(角蛋白、胰蛋白酶)。 最终数据库包含49,878条蛋白质序列。 本数据库适用于基于质谱的蛋白质组学分析流程(如PLGS、Proteome Discoverer)。

提供机构:
Zenodo
创建时间:
2026-04-30
二维码
社区交流群
二维码
科研交流群
商业服务