遇见数据集

Multitask Modeling with Confidence Using Matrix Factorization and Conformal Prediction

收藏
NIAID Data Ecosystem2026-03-11 收录
官方服务:

资源简介:

Multitask prediction of bioactivities is often faced with challenges relating to the sparsity of data and imbalance between different labels. We propose class conditional (Mondrian) conformal predictors using underlying Macau models as a novel approach for large scale bioactivity prediction. This approach handles both high degrees of missing data and label imbalances while still producing high quality predictive models. When applied to ten assay end points from PubChem, the models generated valid models with an efficiency of 74.0–80.1% at the 80% confidence level with similar performance both for the minority and majority class. Also when deleting progressively larger portions of the available data (0–80%) the performance of the models remained robust with only minor deterioration (reduction in efficiency between 5 and 10%). Compared to using Macau without conformal prediction the method presented here significantly improves the performance on imbalanced data sets.

多任务生物活性预测常面临数据稀疏与标签分布失衡的双重挑战。为此,我们提出一种以基础Macau模型为支撑的类别条件(蒙德里安)保角预测器(conformal predictors),将其作为大规模生物活性预测的全新方法。该方法可同时应对高比例缺失数据与标签失衡问题,且仍能构建高质量的预测模型。当将该方法应用于PubChem(公共化学数据库)的10个检测终点任务时,所构建的模型在80%置信水平下的预测效率可达74.0%~80.1%,且对少数类与多数类均能保持相近的预测性能。此外,当逐步移除占比递增的可用数据(0~80%)时,模型性能始终保持稳健,仅出现小幅衰减(效率降低5%~10%)。与未使用保角预测的原生Macau模型相比,本文提出的方法在失衡数据集上的性能得到了显著提升。

创建时间:
2019-04-05
二维码
社区交流群
二维码
科研交流群
商业服务