Automated Detection of Cancer Associated Genes Using a Combined Fuzzy-Rough-Set-Based F-Information and Water Swirl Algorithm of Human Gene Expression Data
收藏资源简介:
This study describes a novel approach to reducing the challenges of highly nonlinear multiclass gene expression values for cancer diagnosis. To build a fruitful system for cancer diagnosis, in this study, we introduced two levels of gene selection such as filtering and embedding for selection of potential genes and the most relevant genes associated with cancer, respectively. The filter procedure was implemented by developing a fuzzy rough set (FR)-based method for redefining the criterion function of f-information (FI) to identify the potential genes without discretizing the continuous gene expression values. The embedded procedure is implemented by means of a water swirl algorithm (WSA), which attempts to optimize the rule set and membership function required to classify samples using a fuzzy-rule-based multiclassification system (FRBMS). Two novel update equations are proposed in WSA, which have better exploration and exploitation abilities while designing a self-learning FRBMS. The efficiency of our new approach was evaluated on 13 multicategory and 9 binary datasets of cancer gene expression. Additionally, the performance of the proposed FRFI-WSA method in designing an FRBMS was compared with existing methods for gene selection and optimization such as genetic algorithm (GA), particle swarm optimization (PSO), and artificial bee colony algorithm (ABC) on all the datasets. In the global cancer map with repeated measurements (GCM_RM) dataset, the FRFI-WSA showed the smallest number of 16 most relevant genes associated with cancer using a minimal number of 26 compact rules with the highest classification accuracy (96.45%). In addition, the statistical validation used in this study revealed that the biological relevance of the most relevant genes associated with cancer and their linguistics detected by the proposed FRFI-WSA approach are better than those in the other methods. The simple interpretable rules with most relevant genes and effectively classified samples suggest that the proposed FRFI-WSA approach is reliable for classification of an individual’s cancer gene expression data with high precision and therefore it could be helpful for clinicians as a clinical decision support system.
本研究提出了一种全新方法,用以缓解癌症诊断中处理高度非线性多分类基因表达值所面临的挑战。为构建高效可用的癌症诊断系统,本研究引入了两级基因选择策略:分别为过滤式基因选择与嵌入式基因选择,以依次筛选潜在致病基因及与癌症关联度最高的相关基因。过滤环节通过构建基于模糊粗糙集(fuzzy rough set, FR)的方法,重新定义了f-信息(f-information, FI)的准则函数,无需对连续型基因表达值进行离散化即可识别潜在基因。嵌入式环节则采用水涡算法(water swirl algorithm, WSA)实现,该算法旨在优化基于模糊规则的多分类系统(fuzzy-rule-based multiclassification system, FRBMS)中样本分类所需的规则集与隶属度函数。本研究针对水涡算法提出了两种全新的更新方程,在设计自学习型FRBMS时具备更优异的探索与开发能力。我们在13个多分类癌症基因表达数据集与9个二分类癌症基因表达数据集上评估了所提方法的效能。此外,将所提出的FRFI-WSA方法在构建FRBMS时的性能,与现有基因选择及优化方法——包括遗传算法(genetic algorithm, GA)、粒子群优化(particle swarm optimization, PSO)及人工蜂群算法(artificial bee colony algorithm, ABC)——在所有数据集上的表现进行了对比。在带有重复测量的全球癌症图谱(global cancer map with repeated measurements, GCM_RM)数据集上,FRFI-WSA方法仅需26条紧凑规则即可筛选出16个与癌症关联度最高的相关基因,同时实现了最高的分类准确率(96.45%)。此外,本研究采用的统计验证结果表明,经FRFI-WSA方法筛选出的与癌症关联度最高的基因及其生物学意义,均优于其他方法所得结果。所提FRFI-WSA方法能够生成兼具高可解释性的规则与高度相关的特征基因,并实现样本的高效分类,这说明该方法在精准分类个体癌症基因表达数据方面具备可靠性能,因此可作为临床决策支持系统为临床医师提供辅助支持。



