Spike-in and real-world proteomics data sets used in publication of PRONE
收藏资源简介:
Spike-in and real-world data sets used in the evaluation study by Arend et al. (see reference), and some utilized in the vignettes of PRONE, an R package designed for preprocessing, normalization, and performance evaluation of normalization methods of proteomics data. Overview of the Data Sets Due to the unavailability of proteomics quantification data from all original publications, data were extracted from alternative sources, which are also listed in the table below. Please refer to the paper's supplementary material and GitHub repository (https://github.com/lisiarend/PRONE.Evaluation) for more comprehensive information on the data sets. A processed metadata file and the protein quantification data file are provided for all data sets. The original quantification data used to generate these two files for each data set are consistently provided in the `original_data` directory of each data set. Data Set Type Quantification Type Raw Data (ID) Quantification Data dS1 UPS1 spike-in (4 levels) LFQ Tabb et al. [1] Välikangas et al. [10] dS2 UPS1 spike-in (6 levels) LFQ Ramus et al. [2] (PXD001819) Graw et al. [11] dS3 E.coli spike-in (5 levels) LFQ Shen et al. [3] (PXD003881) Sticker et al. [12] dS4 E.coli spike-in (2 levels) LFQ Cox et al. [4] (PXD00279) dS5 E.coli spike-in (3 levels) TMT 10-plex (1) Zhu et al. [5] (PXD013277) Phil Wilmarth [13] dS6 yeast spike-in (3 levels) TMT 11-plex (1) O’Connell et al. [6] (PXD007683) Ammar et al. [14] dR1 Osteogenic differentiation of hPCLSCs (4 time points) TMT 6-plex (3) Li et al. [8] (PXD020908) MaxQuant executed in-house dR2 Prospective Ovarian JHU Proteome TMT 10-plex (13) Hu et al. [9] (PDC000110) MaxQuant executed in-house dR3 AROM+ transgenic vs. wild-type mice LFQ Vehmas et al. [7] (PXD002025) dR4 Mycobacterium tuberculosis (healthy, disease vs. treated) TMT 10-plex (2) Schmidt et al. [15] (PXD030883) References [1] D. L. Tabb et al., ‘Repeatability and Reproducibility in Proteomic Identifications by Liquid Chromatography−Tandem Mass Spectrometry’, J. Proteome Res., vol. 9, no. 2, pp. 761–776, Feb. 2010, doi: 10.1021/pr9006365.[2] C. Ramus et al., ‘Spiked proteomic standard dataset for testing label-free quantitative software and statistical methods’, Data Brief, vol. 6, pp. 286–294, Mar. 2016, doi: 10.1016/j.dib.2015.11.063.[3] X. Shen et al., ‘IonStar enables high-precision, low-missing-data proteomics quantification in large biological cohorts’, Proc. Natl. Acad. Sci., vol. 115, no. 21, pp. E4767–E4776, May 2018, doi: 10.1073/pnas.1800541115.[4] J. Cox, M. Y. Hein, C. A. Luber, I. Paron, N. Nagaraj, and M. Mann, ‘Accurate Proteome-wide Label-free Quantification by Delayed Normalization and Maximal Peptide Ratio Extraction, Termed MaxLFQ *’, Mol. Cell. Proteomics, vol. 13, no. 9, pp. 2513–2526, Sep. 2014, doi: 10.1074/mcp.M113.031591.[5] Y. Zhu et al., ‘DEqMS: A Method for Accurate Variance Estimation in Differential Protein Expression Analysis *’, Mol. Cell. Proteomics, vol. 19, no. 6, pp. 1047–1057, Jun. 2020, doi: 10.1074/mcp.TIR119.001646.[6] J. D. O’Connell, J. A. Paulo, J. J. O’Brien, and S. P. Gygi, ‘Proteome-Wide Evaluation of Two Common Protein Quantification Methods’, J. Proteome Res., vol. 17, no. 5, pp. 1934–1942, May 2018, doi: 10.1021/acs.jproteome.8b00016.[7] A. P. Vehmas et al., ‘Liver lipid metabolism is altered by increased circulating estrogen to androgen ratio in male mouse’, J. Proteomics, vol. 133, pp. 66–75, Feb. 2016, doi: 10.1016/j.jprot.2015.12.009.[8] J. Li et al., ‘Dynamic proteomic profiling of human periodontal ligament stem cells during osteogenic differentiation’, Stem Cell Res. Ther., vol. 12, no. 1, p. 98, Feb. 2021, doi: 10.1186/s13287-020-02123-6.[9] Y. Hu et al., ‘Integrated Proteomic and Glycoproteomic Characterization of Human High-Grade Serous Ovarian Carcinoma’, Cell Rep., vol. 33, no. 3, p. 108276, Oct. 2020, doi: 10.1016/j.celrep.2020.108276.[10] T. Välikangas, T. Suomi, and L. L. Elo, ‘A systematic evaluation of normalization methods in quantitative label-free proteomics’, Brief. Bioinform., vol. 19, no. 1, pp. 1–11, Jan. 2018, doi: 10.1093/bib/bbw095.[11] S. Graw et al., ‘proteiNorm – A User-Friendly Tool for Normalization and Analysis of TMT and Label-Free Protein Quantification’, ACS Omega, vol. 5, no. 40, pp. 25625–25633, Oct. 2020, doi: 10.1021/acsomega.0c02564.[12] A. Sticker, L. Goeminne, L. Martens, and L. Clement, ‘Robust Summarization and Inference in Proteome-wide Label-free Quantification’, Mol. Cell. Proteomics, vol. 19, no. 7, pp. 1209–1219, Jul. 2020, doi: 10.1074/mcp.RA119.001624.[13] ‘understanding_IRS’. Accessed: Mar. 07, 2024. [Online]. Available: https://pwilmart.github.io/IRS_normalization/understanding_IRS.html[14] C. Ammar, M. Gruber, G. Csaba, and R. Zimmer, ‘MS-EmpiRe Utilizes Peptide-level Noise Distributions for Ultra-sensitive Detection of Differentially Expressed Proteins[S]’, Mol. Cell. Proteomics, vol. 18, no. 9, pp. 1880–1892, Sep. 2019, doi: 10.1074/mcp.RA119.001509.[15] F. Biadglegne et al., ‘Mycobacterium tuberculosis Affects Protein and Lipid Content of Circulating Exosomes in Infected Patients Depending on Tuberculosis Disease State’, Biomedicines, vol. 10, no. 4, p. 783, Mar. 2022, doi: 10.3390/biomedicines10040783.
本数据集包含Arend等人(见参考文献)的评估研究中使用的掺入标准品(spike-in)数据集与真实世界数据集,其中部分数据集被用于PRONE的示例文档(vignettes)——PRONE是一款用于蛋白质组学数据归一化方法的预处理、归一化及性能评估的R包。 ### 数据集概览 由于无法获取所有原始发表文献中的蛋白质组定量数据,本研究从替代来源提取了数据,这些来源也列于下表。有关数据集的更全面信息,请参阅该论文的补充材料及GitHub仓库(https://github.com/lisiarend/PRONE.Evaluation)。所有数据集均提供了经过处理的元数据文件与蛋白质定量数据文件。用于生成这两份文件的原始定量数据,统一存储于每个数据集的`original_data`目录中。 | 数据集 | 类型 | 定量类型 | 原始数据(ID) | 定量数据 | | ---- | ---- | ---- | ---- | ---- | | dS1 | UPS1掺入标准品(4个梯度) | 无标记定量(Label-Free Quantification, LFQ) | Tabb等人[1] | Välikangas等人[10] | | dS2 | UPS1掺入标准品(6个梯度) | LFQ | Ramus等人[2](PXD001819) | Graw等人[11] | | dS3 | 大肠杆菌掺入标准品(5个梯度) | LFQ | Shen等人[3](PXD003881) | Sticker等人[12] | | dS4 | 大肠杆菌掺入标准品(2个梯度) | LFQ | Cox等人[4](PXD00279) | | | dS5 | 大肠杆菌掺入标准品(3个梯度) | 串联质量标签10-plex(TMT 10-plex (1)) | Zhu等人[5](PXD013277) | Phil Wilmarth[13] | | dS6 | 酵母掺入标准品(3个梯度) | 串联质量标签11-plex(TMT 11-plex (1)) | O’Connell等人[6](PXD007683) | Ammar等人[14] | | dR1 | 人牙周膜干细胞成骨分化(4个时间点) | 串联质量标签6-plex(TMT 6-plex (3)) | Li等人[8](PXD020908) | 内部运行MaxQuant所得 | | dR2 | 前瞻性JHU卵巢蛋白质组 | TMT 10-plex (13) | Hu等人[9](PDC000110) | 内部运行MaxQuant所得 | | dR3 | AROM+转基因小鼠与野生型小鼠对比 | LFQ | Vehmas等人[7](PXD002025) | | | dR4 | 结核分枝杆菌(健康、患病与治疗组) | TMT 10-plex (2) | Schmidt等人[15](PXD030883) | | ### 参考文献 [1] D. L. Tabb et al., ‘Repeatability and Reproducibility in Proteomic Identifications by Liquid Chromatography−Tandem Mass Spectrometry’, J. Proteome Res., vol. 9, no. 2, pp. 761–776, Feb. 2010, doi: 10.1021/pr9006365. [2] C. Ramus et al., ‘Spiked proteomic standard dataset for testing label-free quantitative software and statistical methods’, Data Brief, vol. 6, pp. 286–294, Mar. 2016, doi: 10.1016/j.dib.2015.11.063. [3] X. Shen et al., ‘IonStar enables high-precision, low-missing-data proteomics quantification in large biological cohorts’, Proc. Natl. Acad. Sci., vol. 115, no. 21, pp. E4767–E4776, May 2018, doi: 10.1073/pnas.1800541115. [4] J. Cox, M. Y. Hein, C. A. Luber, I. Paron, N. Nagaraj, and M. Mann, ‘Accurate Proteome-wide Label-free Quantification by Delayed Normalization and Maximal Peptide Ratio Extraction, Termed MaxLFQ *’, Mol. Cell. Proteomics, vol. 13, no. 9, pp. 2513–2526, Sep. 2014, doi: 10.1074/mcp.M113.031591. [5] Y. Zhu et al., ‘DEqMS: A Method for Accurate Variance Estimation in Differential Protein Expression Analysis *’, Mol. Cell. Proteomics, vol. 19, no. 6, pp. 1047–1057, Jun. 2020, doi: 10.1074/mcp.TIR119.001646. [6] J. D. O’Connell, J. A. Paulo, J. J. O’Brien, and S. P. Gygi, ‘Proteome-Wide Evaluation of Two Common Protein Quantification Methods’, J. Proteome Res., vol. 17, no. 5, pp. 1934–1942, May 2018, doi: 10.1021/acs.jproteome.8b00016. [7] A. P. Vehmas et al., ‘Liver lipid metabolism is altered by increased circulating estrogen to androgen ratio in male mouse’, J. Proteomics, vol. 133, pp. 66–75, Feb. 2016, doi: 10.1016/j.jprot.2015.12.009. [8] J. Li et al., ‘Dynamic proteomic profiling of human periodontal ligament stem cells during osteogenic differentiation’, Stem Cell Res. Ther., vol. 12, no. 1, p. 98, Feb. 2021, doi: 10.1186/s13287-020-02123-6. [9] Y. Hu et al., ‘Integrated Proteomic and Glycoproteomic Characterization of Human High-Grade Serous Ovarian Carcinoma’, Cell Rep., vol. 33, no. 3, p. 108276, Oct. 2020, doi: 10.1016/j.celrep.2020.108276. [10] T. Välikangas, T. Suomi, and L. L. Elo, ‘A systematic evaluation of normalization methods in quantitative label-free proteomics’, Brief. Bioinform., vol. 19, no. 1, pp. 1–11, Jan. 2018, doi: 10.1093/bib/bbw095. [11] S. Graw et al., ‘proteiNorm – A User-Friendly Tool for Normalization and Analysis of TMT and Label-Free Protein Quantification’, ACS Omega, vol. 5, no. 40, pp. 25625–25633, Oct. 2020, doi: 10.1021/acsomega.0c02564. [12] A. Sticker, L. Goeminne, L. Martens, and L. Clement, ‘Robust Summarization and Inference in Proteome-wide Label-free Quantification’, Mol. Cell. Proteomics, vol. 19, no. 7, pp. 1209–1219, Jul. 2020, doi: 10.1074/mcp.RA119.001624. [13] ‘understanding_IRS’. Accessed: Mar. 07, 2024. [Online]. Available: https://pwilmart.github.io/IRS_normalization/understanding_IRS.html [14] C. Ammar, M. Gruber, G. Csaba, and R. Zimmer, ‘MS-EmpiRe Utilizes Peptide-level Noise Distributions for Ultra-sensitive Detection of Differentially Expressed Proteins[S]’, Mol. Cell. Proteomics, vol. 18, no. 9, pp. 1880–1892, Sep. 2019, doi: 10.1074/mcp.RA119.001509. [15] F. Biadglegne et al., ‘Mycobacterium tuberculosis Affects Protein and Lipid Content of Circulating Exosomes in Infected Patients Depending on Tuberculosis Disease State’, Biomedicines, vol. 10, no. 4, p. 783, Mar. 2022, doi: 10.3390/biomedicines10040783.



