遇见数据集

个人敏感信息识别系统的数据集

收藏
官方服务:

资源简介:

本数据集用于个人敏感信息识别系统的功能和性能测试;该数据集包含52个场景,涵盖15种类别,5种模态,总计5720万个人信息数据。其中,5种模态数据分别包括个人诚信声明文档、用户身份照片、用户视频数据和用户图形数据,用户的个人数据不涉及个人隐私,均采用模拟数据的方式进行成成,用户信息均进行脱敏操作,其中视频涉及个人信息的内容进行了删减。数据集主目录下中共包含group1-10和multimodal, multimodal文件中包含audio_file、identity_cards、statements_docx、svg_files;statements_docx中包含500个DOCX格式文件;identity_cards中包含500个JPG格式文件;audio_files中包含500个WAV格式文件;svg_files中包含500个SVG格式文件。数据集一共12.5GB,未压缩,满足数据集提交数据量的要求。

This dataset is designed for functional and performance testing of personal sensitive information recognition systems. It consists of 52 scenarios covering 15 categories and 5 modalities, with a total of 57.2 million pieces of personal information data. The 5 modalities include personal integrity statement documents, user identity photos, user video data, user graphic data, and audio data. All personal data in this dataset does not involve real personal privacy, and is generated using simulated data; all user information has been desensitized, and content containing personal information in the video data has been removed. The main directory of the dataset contains 10 groups (group1 to group10) and a multimodal folder. The multimodal folder includes four subfolders: audio_files, identity_cards, statements_docx, and svg_files. Specifically, the statements_docx folder contains 500 DOCX-format files, the identity_cards folder contains 500 JPG-format files, the audio_files folder contains 500 WAV-format files, and the svg_files folder contains 500 SVG-format files. The total size of the uncompressed dataset is 12.5 GB, which meets the data volume requirements for dataset submission.

搜集汇总
数据集介绍
个人敏感信息识别系统的数据集 数据集图片
背景与挑战
背景概述
该数据集专为个人敏感信息识别系统的功能与性能测试设计,涵盖52个场景、15种类别和5种模态,包括文档、照片、视频等多种格式,总计5720万条经过脱敏处理的模拟个人信息数据。数据集总大小为12.5GB,未压缩,满足相关数据量要求。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务