遇见数据集

Privacy-Utility Trade-offs in Differentially Private Synthetic Tabular Data: a reproducible benchmark on the UCI Heart Disease (Cleveland) dataset

收藏
Zenodo2026-09-30 更新2026-10-01 收录
官方服务:

资源简介:

A reproducible benchmark measuring how the utility of differentially private (DP) synthetic tabular data degrades as the privacy budget tightens, and separating the error introduced by the synthesizer from that introduced by the downstream DP estimator. Contents heart_benchmark_results.csv — tidy results, 2,444 rows x 15 columns. One row per (synthesizer, epsilon, replicate, method, statistic), carrying true_value (ground truth computed on the original data), dp_estimate, and non_dp_estimate (a non-private control computed on the same synthetic data). Comparing synth_abs_error with abs_error separates synthesis loss from inference loss. metadata_dictionary.csv — column-by-column data dictionary, including notes on noise layering, provenance, and design conditioning. reproduce_heart_benchmark.R — annotated end-to-end reproduction script, runtime approximately 40 seconds. data/heart_cleveland_clean.csv — cleaned input, 297 complete records. figures/heart_utility_tradeoff.png — four-panel privacy-utility trade-off figure. synthesis_failures.csv — replicate-level synthesis failures with originating error messages. MANUSCRIPT.md and the compiled Data Descriptor manuscript (DOCX). Method. Four DP synthesizers (Gaussian copula, histogram marginals, PATE, Gaussian mixture) across five privacy budgets (epsilon = 0.1, 0.5, 1, 2, 8) and ten Monte Carlo replicates, scored with differentially private descriptive statistics, two hypothesis tests, and linear regression. Master seed 20,260,929, with per-cell seeds derived deterministically so that any single cell reproduces in isolation. Principal finding. The commonly specified regression cholesterol ~ age + resting_blood_pressure is so ill-conditioned on this data that its differentially private coefficients are pure noise at every privacy budget examined: the smallest eigenvalue of the design matrix is 4.09, against 22,019.2 for a publicly centred, intercept-free alternative. This is an 18,691-fold difference in DP sensitivity arising purely from parameterisation, with no change to the data, the response bounds, or the privacy budget. Both specifications are released so that the contrast is auditable. The centred specification yields the expected monotone trade-off, with coefficient MSE falling from 1,695 at epsilon = 0.1 to 1.41 at epsilon = 8. Negative results. The PATE synthesizer shows no privacy-utility trade-off at all: MSE is constant at 23.76 with an interquartile range of 0.02 across an 80-fold change in epsilon. The Gaussian mixture synthesizer releases only 6 of 15 columns, cannot support the categorical tasks, and produces a non-monotone error curve. The copula and histogram-marginal synthesizers fail outright at epsilon = 0.1 in 4 of 50 replicate cells, where the DP-perturbed correlation matrix ceases to be positive definite and the multivariate normal sampler rejects it. Provenance of the source data. Derived from the Cleveland subset of the UCI Heart Disease dataset (UCI Machine Learning Repository, dataset id 45, https://archive.ics.uci.edu/dataset/45/heart+disease). The raw repository release is not redistributed here; only the cleaned and derived files listed above are included. Six records with missing values in the attributes ca (4 records) and thal (2 records) were removed rather than imputed, retaining 297 complete records. Declaration of interests. The creator of this deposit is also the author and maintainer of both DPSynth and DPrivStats, the differentially private software packages evaluated in this benchmark. This is a self-evaluation, declared in full in the associated manuscript (Competing Interests). The benchmark reports unfavourable results for the creator's own packages, including the synthesis failures and structural limitations summarised above. All four synthesizers were run under an identical protocol, using the packages exactly as published on CRAN without modification, and the complete results table is released so that every comparison can be independently re-derived.

提供机构:
Bijalwan, Mukul
创建时间:
2026-09-29
二维码
社区交流群
二维码
科研交流群
商业服务