遇见数据集

Performance for lung data.

收藏
Figshare2024-05-10 更新2026-04-28 收录
官方服务:

资源简介:

In recent years, researchers have proven the effectiveness and speediness of machine learning-based cancer diagnosis models. However, it is difficult to explain the results generated by machine learning models, especially ones that utilized complex high-dimensional data like RNA sequencing data. In this study, we propose the binarilization technique as a novel way to treat RNA sequencing data and used it to construct explainable cancer prediction models. We tested our proposed data processing technique on five different models, namely neural network, random forest, xgboost, support vector machine, and decision tree, using four cancer datasets collected from the National Cancer Institute Genomic Data Commons. Since our datasets are imbalanced, we evaluated the performance of all models using metrics designed for imbalance performance like geometric mean, Matthews correlation coefficient, F-Measure, and area under the receiver operating characteristic curve. Our approach showed comparative performance while relying on less features. Additionally, we demonstrated that data binarilization offers higher explainability by revealing how each feature affects the prediction. These results demonstrate the potential of data binarilization technique in improving the performance and explainability of RNA sequencing based cancer prediction models.

近年来,研究者已证实基于机器学习的癌症诊断模型兼具有效性与高效性。然而,机器学习模型所生成的结果往往难以解释,尤其是那些使用RNA测序(RNA sequencing)这类复杂高维数据的模型。本研究提出二值化技术作为处理RNA测序数据的全新方案,并以此构建可解释的癌症预测模型。我们将所提出的数据处理技术应用于五种不同模型开展测试,分别为神经网络(neural network)、随机森林(random forest)、极端梯度提升(xgboost)、支持向量机(support vector machine)以及决策树(decision tree),所用的4个癌症数据集均采集自美国国家癌症研究所基因组数据公共库(National Cancer Institute Genomic Data Commons)。由于本研究的数据集存在类别不平衡问题,我们采用针对不平衡分类性能设计的评估指标对所有模型的性能进行评测,包括几何均值(geometric mean)、马修斯相关系数(Matthews correlation coefficient)、F测度(F-Measure)以及受试者工作特征曲线下面积(area under the receiver operating characteristic curve)。我们的方法在依赖更少特征的前提下,取得了相当的模型性能。此外,我们通过揭示每个特征对预测结果的影响机制,证实了数据二值化技术具备更高的可解释性。上述结果证明,数据二值化技术在提升基于RNA测序数据的癌症预测模型的性能与可解释性方面具有可观的应用潜力。

创建时间:
2024-05-10
二维码
社区交流群
二维码
科研交流群
商业服务