gretel-synthetic-pii-personas-v1
收藏资源简介:
该数据集包含使用隐私保护技术生成的合成人物数据,适用于测试数据隐私工具、开发姓名识别模型和其他与身份相关的机器学习任务。数据集包括多样化的个人身份信息(PII),统计信息显示总样本数为991,508,涵盖205个独特国家、404,175个独特公司、86,542个独特名字和241,682个独特姓氏。数据集展示了国家、地区和名字的分布情况,并提供了一个完整的数据记录示例。数据集适用于测试数据隐私工具、开发姓名识别模型等,但数据是合成的,不应用于对真实个体或群体的推断。
This dataset contains synthetic personal data generated via privacy-preserving technologies, which is applicable for testing data privacy tools, developing name recognition models and other identity-related machine learning tasks. It includes diverse Personally Identifiable Information (PII). Statistical results show that the total number of samples is 991,508, covering 205 unique countries, 404,175 unique companies, 86,542 unique given names and 241,682 unique family names. The dataset presents the distribution of countries, regions and names, and provides a complete data record example. Although the dataset is suitable for testing data privacy tools, developing name recognition models and other similar tasks, the data is synthetic and should not be used for inferences about real individuals or groups.




