遇见数据集

Accurate and fast feature selection workflow for high-dimensional omics data

收藏
Figshare2017-12-21 更新2026-04-29 收录
官方服务:

资源简介:

We are moving into the age of ‘Big Data’ in biomedical research and bioinformatics. This trend could be encapsulated in this simple formula: D = S * F, where the volume of data generated (D) increases in both dimensions: the number of samples (S) and the number of sample features (F). Frequently, a typical omics classification includes redundant and irrelevant features (e.g. genes or proteins) that can result in long computation times; decrease of the model performance and the selection of suboptimal features (genes and proteins) after the classification/regression step. Multiple algorithms and reviews has been published to describe all the existing methods for feature selection, their strengths and weakness. However, the selection of the correct FS algorithm and strategy constitutes an enormous challenge. Despite the number and diversity of algorithms available, the proper choice of an approach for facing a specific problem often falls in a ‘grey zone’. In this study, we select a subset of FS methods to develop an efficient workflow and an R package for bioinformatics machine learning problems. We cover relevant issues concerning FS, ranging from domain’s problems to algorithm solutions and computational tools. Finally, we use seven different proteomics and gene expression datasets to evaluate the workflow and guide the FS process.

我们正步入生物医学研究与生物信息学领域的“大数据”时代。这一趋势可由一个简洁公式概括:D = S × F,其中生成数据的总量(D)随两个维度增长:样本数量(S)与样本特征数量(F)。通常,典型的组学分类任务会包含冗余且无关的特征(如基因或蛋白质),这会导致计算时长增加、模型性能下降,且在分类/回归步骤后筛选出的特征(基因与蛋白质)并非最优。已有大量算法与综述文献对现有特征选择(Feature Selection, FS)方法及其优缺点进行了阐述。然而,选择合适的特征选择算法与策略仍是一项极具挑战性的任务。尽管现有算法数量繁多、类型多样,但针对特定问题选择恰当的解决方案往往处于“灰色地带”。本研究选取了部分特征选择方法,旨在开发一套适用于生物信息学机器学习任务的高效工作流与R语言工具包。本研究涵盖了与特征选择相关的各类议题,从领域专属问题到算法解决方案与计算工具均有涉及。最后,我们使用7组不同的蛋白质组学与基因表达数据集对该工作流进行评估,并为特征选择流程提供指导。

创建时间:
2017-12-21
二维码
社区交流群
二维码
科研交流群
商业服务