遇见数据集

Read Me from Constraining classifiers in molecular analysis: invariance and robustness

收藏
DataCite Commons2020-08-26 更新2024-07-28 收录
官方服务:

资源简介:

Analysing molecular profiles requires the selection of classification models that can cope with the high dimensionality and variability of this data. Also, improper reference point choice and scaling pose additional challenges. Often model selection is somewhat guided by <i>ad hoc</i> simulations rather than by sophisticated considerations on the properties of a categorization model. Here, we derive and report four linked linear concept classes/models with distinct invariance properties for high-dimensional molecular classification. We can further show that these concept classes also form a half-order of complexity classes in terms of Vapnik–Chervonenkis dimensions, which also implies increased generalization abilities. We implemented support vector machines with these properties. Surprisingly, we were able to attain comparable or even superior generalization abilities to the standard linear one on the 27 investigated RNA-Seq and microarray datasets. Our results indicate that <i>a priori</i> chosen invariant models can replace <i>ad hoc</i> robustness analysis by interpretable and theoretically guaranteed properties in molecular categorization.

分析分子谱(molecular profiles)时,需选择可适配该数据高维度与变异性的分类模型。此外,参考点选择不当与特征缩放亦会带来额外挑战。当前模型选择往往在一定程度上由权宜性(ad hoc)模拟主导,而非基于对分类模型属性的严谨考量。本研究推导并报道了四类具备不同不变性特性的关联线性概念类/模型,适用于高维分子分类任务。进一步研究表明,这些概念类基于VC维(Vapnik–Chervonenkis dimension)可构成复杂度类的半序关系,这同时意味着其泛化能力得到提升。我们针对上述特性实现了支持向量机(support vector machines)。令人意外的是,在27个被研究的RNA测序(RNA-Seq)与基因芯片(microarray)数据集上,我们的模型获得了与标准线性模型相当甚至更优的泛化能力。本研究结果显示,在分子分类任务中,预先选定(a priori)的不变性模型可通过具备可解释性与理论保障的特性,替代权宜性的鲁棒性分析。

提供机构:
The Royal Society
创建时间:
2020-01-21
二维码
社区交流群
二维码
科研交流群
商业服务