遇见数据集

Regularized Optimal Transport of Covariates and Outcomes in Data Recoding

收藏
DataCite Commons2020-08-25 更新2024-07-28 收录
官方服务:

资源简介:

When databases are constructed from heterogeneous sources, it is not unusual that different encodings are used for the same outcome. In such case, it is necessary to recode the outcome variable before merging two databases. The method proposed for the recoding is an application of optimal transportation where we search for a bijective mapping between the distributions of such variable in two databases. In this article, we build upon the work by Garés et al., where they transport the distributions of categorical outcomes assuming that they are distributed equally in the two databases. Here, we extend the scope of the model to treat all the situations where the covariates explain the outcomes similarly in the two databases. In particular, we do not require that the outcomes be distributed equally. For this, we propose a model where joint distributions of outcomes and covariates are transported. We also propose to enrich the model by relaxing the constraints on marginal distributions and adding an <i>L</i><sup>1</sup> regularization term. The performances of the models are evaluated in a simulation study, and they are applied to a real dataset. The code used in the computational assessment and in the simulation of test cases is publicly available on Github repository: https://github.com/otrecoding/OTRecod.jl.

当从异构数据源构建数据库时,同一结果变量采用不同编码的情况并不罕见。在此类场景下,合并两个数据库前需对结果变量进行重编码处理。本次提出的重编码方法基于最优传输(optimal transportation)技术,旨在寻找两个数据库中该变量分布间的双射映射关系。本文基于Garés等人的研究工作展开:该团队在假设两个数据库中分类结果分布一致的前提下,完成了分类结果的分布传输。而本文则拓展了该模型的适用范围,可处理两个数据库中协变量(covariates)对结果的解释逻辑一致的全部场景,尤为关键的是,本文不再要求结果变量的分布保持一致。为此,本文提出一种对结果变量与协变量的联合分布进行传输的模型。此外,本文还通过放松对边缘分布的约束并添加L¹正则项(L¹ regularization term)对模型进行优化扩展。本文通过仿真实验评估了所提模型的性能,并将其应用于真实数据集。用于计算评估与测试用例仿真的代码已公开至Github仓库:https://github.com/otrecoding/OTRecod.jl。

提供机构:
Taylor & Francis
创建时间:
2020-06-01
二维码
社区交流群
二维码
科研交流群
商业服务