遇见数据集

K-RANK: AN EVOLUTION OF Y-RANK FOR MULTIPLE SOLUTIONS PROBLEM

收藏
NIAID Data Ecosystem2026-03-11 收录
官方服务:

资源简介:

ABSTRACT Y-rank can present faults when dealing with non-linear problems. A methodology is proposed to improve the selection of data in situations where y-rank is fragile. The proposed alternative, called k-rank, consists of splitting the data set into clusters using the k-means algorithm, and then apply y-rank to the generated clusters. Models were calibrated and tested with subsets split by y-rank and k-rank. For the Heating Tank case study, in 59% of the simulations, models calibrated with k-rank subsets achieved better results. For the Propylene / Propane Separation Unit case, when dealing with a small number of sample points, the y-rank models had errors almost three times higher than the k-rank models for the test subset, meaning that the fitted model could not deal properly with new unseen data. The proposed methodology was successful in splitting the data, especially in cases with a limited amount of samples.

摘要:Y-rank在处理非线性问题时存在缺陷。针对Y-rank表现欠佳的场景,本文提出一种优化数据筛选流程的方法。该替代方案命名为k-rank,其具体实现为:先通过k-means聚类算法将数据集划分为若干簇,随后对生成的各簇应用Y-rank方法。研究人员分别使用Y-rank与k-rank划分得到的子集对模型进行校准与测试。在加热罐(Heating Tank)案例研究中,59%的模拟结果显示,基于k-rank子集校准的模型表现更优。在丙烯/丙烷分离装置(Propylene / Propane Separation Unit)案例中,当样本点数量较少时,测试子集上的Y-rank模型误差几乎是k-rank模型的三倍,表明拟合得到的模型无法妥善处理全新的未见数据。所提出的方法在数据划分任务中效果优异,尤其适用于样本量有限的场景。

创建时间:
2019-03-01
二维码
社区交流群
二维码
科研交流群
商业服务