脱敏效果评测系统的模拟数据集
收藏资源简介:
在脱敏效果评测系统中,样本数据主要用于为评测提供比对依据。样本数据基于个人敏感信息按需脱敏工具集模拟数据集构建。对于表格和文本模态数据,表格模态数据个人敏感信息按需脱敏工具集模拟数据集中的52个典型应用场景数据集构建,对于图像、视频、音频、图形模态数据针对非结构化数据,图像、视频、音频模态与个人敏感信息按需脱敏工具集模拟数据集中图像、视频、音频模态的图像、视频、音频、图形各模态1000条数据保持一致。样例数据用于验证脱敏算法在不同模态数据上的表现。样例数据的生成基于6000万人的个人信息数据,并结合文本、表格和非结构化数据的特定脱敏算法,构造多模态脱敏样例。1)对于文本和表格模态数据,基于场景数据集,结合多种脱敏算法,如随机置换、尾部截断等,生成不同版本的脱敏样例数据。2)对于如图像、视频、音频和图形模态数据,通过特定的脱敏技术生成脱敏样例。具体方法,针对图像数据进行模糊处理或图像加扰;针对视频数据应用视频模糊;针对音频数据进行频率变换;针对轨迹数据进行随机扰动。
In the de-identification effect evaluation system, sample data primarily serves as a comparison basis for performance evaluations. The sample data is constructed from the simulated dataset generated by the on-demand personal sensitive information de-identification toolset. For tabular and text modal data, the tabular dataset is built using 52 typical application scenario datasets from the simulated dataset of the on-demand personal sensitive information de-identification toolset. For unstructured data including image, video, audio and graphic modalities, the data for image, video and audio modalities are consistent with the 1000 samples per modality (image, video, audio and graphic) included in the aforementioned simulated dataset. The sample data is used to validate the performance of de-identification algorithms across various modal data. The sample data is generated based on the personal information data of 60 million people, combined with specific de-identification algorithms tailored for text, tabular and unstructured data to create multi-modal de-identification samples: 1) For text and tabular modal data, multiple versions of de-identified sample data are generated by combining the scenario datasets with various de-identification algorithms such as random permutation, tail truncation and others. 2) For modal data including image, video, audio and graphic, de-identified samples are generated via specific de-identification techniques. The specific methods are as follows: applying blurring or image scrambling for image data; using video blurring for video data; performing frequency transformation for audio data; and conducting random perturbation for trajectory data.




