遇见数据集

Databases of human SARS-CoV-2 antibody peptides for bottom-up proteomics

收藏
Zenodo2024-09-17 更新2026-05-26 收录
官方服务:

资源简介:

Bottom-up proteomics approaches rely on database searches that compare experimental values of peptides to theoretical values derived from protein sequences in a database. While the human body can produce millions of distinct antibodies, current databases for human antibodies such as UniProtKB are limited to only 1095 sequences (as of 2024 January). This limitation may hinder the identification of new antibodies using bottom-up proteomics. Therefore, extending the databases is an important task for discovering new antibodies. Herein, we adopted extensive collection of antibody sequences from Observed Antibody Space for conducting efficient database searches in publicly available proteomics data with a focus on the SARS-CoV-2 disease. Thirty million heavy antibody sequences from 146 SARS-CoV-2 patients in the Observed Antibody Space were in silico digested to obtain 18 million unique peptides. These peptides were then used to create six databases (DB1-DB6) for bottom-up proteomics. We used those databases for searching antibody peptides in publicly available SARS-CoV-2 human plasma samples in the Proteomics Identification Database (PRIDE), and we consistently found new antibody peptides in those samples. The database searching task was done by using Fragpipe softwares. Table 1. Information of databases. In addition to human SARS-CoV-2 antibody peptides, every database also contains human protein sequences from UniProt database and contaminants from cRAP database. File Database Number of human SARS-CoV-2 antibody peptides DB1.fasta DB1 100 DB2.fasta DB2 1,000 DB3.fasta DB3 10,000 DB4.fasta DB4 100,000 DB5.fasta DB5 1,000,000 DB6.fasta DB6 10,000,000

提供机构:
Zenodo
创建时间:
2024-01-24
二维码
社区交流群
二维码
科研交流群
商业服务