遇见数据集

Anonymized Datasets and LLM Imputation/Predition for Adult Income, German Credit, Diabetes Readmission, Employment

收藏
Zenodo2026-02-04 更新2026-05-26 收录
官方服务:

资源简介:

Data repository for the GitHub: https://github.com/luckyos-code/user-driven-privacy which is source to the Paper: Learning from Anonymized and Incomplete Tabular Data Please cite this paper as reference. Content Information Original datasets used: German Credit Adult Census Income Diabetes Readmission (equal split, binary classes) Employment (2018, California) This data repository contains the following folders: 'data/': original and anonymized data for four datasets sorted into subfolders by dataset name at the toplevel of each dataset folder we find a train.csv and test.csv with original data in 80/20 split and the following subfolders: 'generalization/' subfolder contains the anonymized data that again is sorted into subfolders for the different privacy distributions, with train and test data for each 'forced_generalization/' similar structure to `generalization/` but contains the prepared data under forced generalization method on the anonymized datasets, contains the whole dataset (no split) which will then be joined into the respective train and test parts by record ids in the experiments Note: specialization data and the resulting datasets of other method are not available here as they are created and handled in-memory only 'llm_evaluation/': results of llm-based methods on the anonymized datasets at top level we find subfolders per privacy distribution that each contains the output data for all datasets anonymized with this distribution: six files per dataset with results for train and test data separately and split into the anonymized values together with the respectively imputed values for a dataset, the resulting imputed dataset, and the direct label prediction outcomes on the given dataset

提供机构:
Zenodo
创建时间:
2026-01-31
二维码
社区交流群
二维码
科研交流群
商业服务