C-glutamicum
收藏资源简介:
C. Glutamicum数据集是一个蛋白质组学质谱数据集,专注于谷氨酸棒状杆菌(Corynebacterium glutamicum)的Top-down蛋白质分析。该数据集源自TopRepo,一个包含超过1200万MS/MS谱图、覆盖12个物种的Top-down谱图库。数据通过标准化的分析流程生成:原始质谱文件经msconvert转换为质心mzML文件,再使用TopFD进行谱图去卷积,生成msalign文件和蛋白型特征文件,最后通过TopPIC搜索相应的蛋白质组序列数据库进行谱图鉴定。本数据集利用该流程产生的mzML文件、msalign文件、特征文件和谱图鉴定结果(TSV文件),通过专用Python脚本进一步处理,生成了包含综合谱图信息的结构化数据。数据集包含三个核心文件:1) 元数据表(toprepo_c_glutamicum_meta_table_v1.2.0.tsv),提供实验级别的信息,包括数据集标识符、物种信息、仪器元数据、解离方法以及谱图、蛋白质和蛋白型的计数统计;2) 谱图表(toprepo_c_glutamicum_spectrum_table_ms2_v1.2.0.tsv),包含MS/MS谱图级别的分析数据,如扫描标识符、前体离子信息(质量、电荷)、碎片测量值、蛋白型注释、蛋白质鉴定结果以及统计置信度指标;3) 注释的.msalign文件。该数据集适用于蛋白质鉴定、蛋白型表征、质谱数据分析方法开发等生物信息学和蛋白质组学研究任务。
The C. Glutamicum dataset is a proteomic mass spectrometry dataset focused on top-down protein analysis of Corynebacterium glutamicum. It originates from TopRepo, a top-down spectral library containing over 12 million MS/MS spectra covering 12 species. The data is generated through a standardized analytical pipeline: raw mass spectrometry files are converted to centroid mzML files using msconvert, then deconvoluted with TopFD to produce msalign files and proteoform feature files, followed by spectral identification via TopPIC against a corresponding proteome sequence database. This dataset leverages the resulting mzML files, msalign files, feature files, and spectral identification results (TSV files), which are further processed using a dedicated Python script to generate structured data with comprehensive spectral information. The dataset includes three core files: 1) a metadata table (toprepo_c_glutamicum_meta_table_v1.2.0.tsv) providing experiment-level information such as dataset identifiers, species details, instrument metadata, dissociation methods, and counts of spectra, proteins, and proteoforms; 2) a spectrum table (toprepo_c_glutamicum_spectrum_table_ms2_v1.2.0.tsv) containing MS/MS spectrum-level analytical data, including scan identifiers, precursor ion information (mass, charge), fragment measurements, proteoform annotations, protein identification results, and statistical confidence metrics; and 3) annotated .msalign files. It is suitable for bioinformatics and proteomics research tasks such as protein identification, proteoform characterization, and mass spectrometry data analysis method development.
数据集概览:TopRepo / C. glutamicum
名称:C. Glutamicum Dataset
许可证:Apache-2.0
语言:英文
数据集背景
TopRepo 是一个自上而下的质谱谱图库,包含来自 12 个物种的超过 1200 万个 MS/MS 谱图。数据处理流程包括:将原始文件转换为中心化 mzML 文件,使用 TopFD 进行解卷积生成 msalign 文件和蛋白质形态特征文件,再通过 TopPIC 搜索蛋白质组数据库进行谱图鉴定,结果存储于 TSV 文件中。
数据集结构
针对物种 Corynebacterium glutamicum,数据集包含以下文件:
-
元数据表 (
toprepo_c_glutamicum_meta_table_v1.2.0.tsv)
包含实验级别的元数据,如数据集标识符、物种信息、仪器元数据、解离方法、谱图/蛋白质/蛋白质形态计数。 -
谱图表 (
toprepo_c_glutamicum_spectrum_table_ms2_v1.2.0.tsv)
包含谱图级别的 MS/MS 分析数据,如扫描标识符、前体离子信息、碎裂测量值、蛋白质形态注释、蛋白质鉴定结果和统计置信度指标。 -
带注释的 .msalign 文件
数据来源
- 在线谱图库:https://toprepo.org/
- 论文:https://www.biorxiv.org/content/10.64898/2026.02.20.707032v1
- GitHub 仓库:https://github.com/toppic-suite/toprepo
联系方式
如有疑问,请联系 Tulane 大学 TopRepo 团队:xwliu@tulane.edu 或 kli7@tulane.edu




