Supporting data for "A scalable software solution for anonymizing high-dimensional biomedical data"
收藏资源简介:
Data anonymization is an important building block for ensuring privacy and fosters the re-use of data. However, transforming the data in a way it preserves the privacy of subjects while maintaining a high degree of data quality is challenging and particularly difficult when processing complex datasets that contain a high number of attributes. In this paper we present how we extended the open-source software ARX to improve its support for high-dimensional, biomedical datasets. <br>For improving ARXs capability to find optimal transformations when processing high-dimensional data, we implement two novel search algorithms. The first one is a greedy top-down approach and is oriented on a formally implemented bottom-up search. The second is based on a genetic algorithm. We evaluated the algorithms with different datasets, transformation methods and privacy models. The novel algorithms mostly outperformed the previously implemented bottom-up search. Additionally, we extended the graphical user interface to provide a high degree of usability and performance when working with high-dimensional datasets. <br>With our additions we have significantly enhanced ARXs ability to handle high-dimensional data in terms of processing performance as well as usability and thus can further facilitate data sharing.



