Incomplete big datasets for parallel fractional hot-deck imputation

Name: Incomplete big datasets for parallel fractional hot-deck imputation
Creator: IEEE DataPort
Published: 2021-07-02 20:24:47
License: 暂无描述

DataCite Commons2021-07-02 更新2025-04-16 收录

下载链接：

https://ieee-dataport.org/open-access/incomplete-big-datasets-parallel-fractional-hot-deck-imputation

下载链接

链接失效反馈

官方服务：

资源简介：

Parallel fractional hot-deck imputation (P-FHDI) is a general-purpose, assumption-free tool for handling item nonresponse in big incomplete data by combining the theory of FHDI and parallel computing. FHDI cures multivariate missing data by filling each missing unit with multiple observed values (thus, hot-deck) without resorting to distributional assumptions. P-FHDI can tackle big incomplete data with millions of instances (big-n) or 10, 000 variables (big-p). However, handling ultra incomplete data (i.e., concurrently big-n and big-p) with tremendous instances and high dimensionality has posed challenges to P-FHDI due to excessive memory requirement and execution time. We developed the ultra data-oriented P-FHDI (named as UP-FHDI) capable of curing ultra incomplete data. In addition to the parallel Jackknife method, UP-FHDI enables a computationally efficient ultra data-oriented variance estimation by using parallel linearization techniques. The object of this repository is to illustrate the use of UP-FHDI with different types of big incomplete data.

提供机构：

IEEE DataPort

创建时间：

2021-07-02

5,000+

优质数据集

54 个

任务类型

进入经典数据集