Zebrafish
收藏资源简介:
该数据集是TopRepo(一个自上而下质谱图库)中针对斑马鱼物种的子集。TopRepo整体包含来自12个物种的超过1200万个MS/MS质谱图。本数据集通过一套标准分析流程生成:原始质谱文件经msconvert转换为centroided mzML文件,再通过TopFD进行谱图去卷积,生成msalign文件和proteoform特征文件。随后,使用TopPIC将msalign文件与斑马鱼蛋白质组序列数据库比对,实现谱图鉴定,结果存储在TSV文件中。基于此流程,Python脚本进一步生成了包含全面谱图信息的TSV文件、注释后的msalign文件以及注释后的mgf文件。数据集核心由两个TSV文件构成:元数据表(包含实验级别元数据如数据集标识符、物种信息、仪器元数据等)和质谱图表(包含谱图级别MS/MS分析数据如扫描标识符、前体离子信息、碎片离子测量值等)。此外,数据集还包含一个注释后的.msalign文件。该数据集适用于蛋白质组学、生物信息学领域的研究,特别是基于质谱的蛋白质鉴定、蛋白质形态分析、谱图库构建及相关算法开发与验证任务。
This dataset is a subset of TopRepo (a top-down mass spectral library) for the Zebrafish species. TopRepo overall contains over 12 million MS/MS spectra from 12 species. The dataset is generated through a standard analysis pipeline: raw mass spectrometry files are converted to centroided mzML files via msconvert, then deconvoluted using TopFD to produce msalign files and proteoform feature files. Subsequently, TopPIC is used to align the msalign files with the corresponding Zebrafish proteome sequence database for spectral identification, with results stored in TSV files. Based on this pipeline, Python scripts further generate TSV files with comprehensive spectral information, annotated msalign files, and annotated mgf files. The core of the dataset consists of two TSV files: a metadata table (containing experiment-level metadata such as dataset identifiers, species information, instrument metadata, etc.) and a mass spectrum table (containing spectrum-level MS/MS analysis data such as scan identifiers, precursor ion information, fragment ion measurements, etc.). Additionally, the dataset includes an annotated .msalign file. It is suitable for research in proteomics and bioinformatics, particularly for protein identification, proteoform analysis, spectral library construction, and related algorithm development and validation based on mass spectrometry (especially top-down mass spectrometry and LC-MS).
数据集概述
- 数据集名称:Zebrafish Dataset (TopRepo 子集)
- 语言:英语
- 许可证:Apache-2.0
- 标签:蛋白质组学、生物信息学、质谱分析、液相色谱-质谱联用、谱库、蛋白质、蛋白质鉴定
- 来源:TopRepo,一个包含12个物种、超过1200万张串联质谱(MS/MS)光谱的顶级质谱库。
数据来源与处理流程
- 原始数据转换:每个质谱原始文件(.raw)通过
msconvert转换为中心化的 mzML 文件。 - 谱图解卷积:使用
TopFD将 mzML 文件中的光谱解卷积为一个或多个msalign文件及蛋白质形式特征文件。 - 谱图鉴定:使用
TopPIC将msalign文件与其对应的蛋白质组序列数据库进行比对搜索,鉴定结果存储在 TSV 文件中。 - 结果生成:基于上述流程产生的 mzML、msalign、特征文件和鉴定结果(TSV),通过 Python 脚本生成包含全面光谱信息的 TSV 文件、注释后的
msalign文件以及注释后的 mgf 文件。
数据集结构
该数据集包含针对斑马鱼(Zebrafish)物种的以下文件:
- 元数据表 (
toprepo_zebrafish_meta_table_v1.2.0.tsv):包含实验级别元数据。- 数据集标识符
- 物种信息
- 仪器元数据
- 解离方法
- 光谱、蛋白质和蛋白质形式的数量统计
- 光谱表 (
toprepo_zebrafish_spectrum_table_ms2_v1.2.0.tsv):包含光谱级别(MS/MS)的分析数据。- 扫描标识符
- 前体离子信息
- 碎裂测量数据
- 蛋白质形式注释
- 蛋白质鉴定结果
- 统计置信度指标
- 注释后的 .msalign 文件
相关资源
- 仓库主页:https://toprepo.org/
- 相关论文:https://www.biorxiv.org/content/10.64898/2026.02.20.707032v1
- GitHub 仓库:https://github.com/toppic-suite/toprepo
联系方式
- 如有疑问,可联系 TopRepo 团队:xwliu@tulane.edu 或 kli7@tulane.edu。
版权信息
- 版权所有 (c) 2025 - 2026, 杜兰大学 (Tulane University)。




