遇见数据集

cicero-im/modified

收藏
Hugging Face2025-03-17 更新2025-04-26 收录
官方服务:

资源简介:

该数据集包含葡萄牙语的匿名化示例,由原始文本、遮蔽后的文本、识别出的个人信息识别(PII)实体以及匿名化过程中可能引入的数据污染信息组成。

This dataset contains anonymization examples in Portuguese, consisting of the original text, masked text, identified PII entities, and potential data pollution introduced during the anonymization process.

提供机构:
cicero-im
二维码
社区交流群
二维码
科研交流群
商业服务