cs419-data
收藏资源简介:
该数据集是一个用于隐私敏感信息标注的文本数据集,包含训练集(54,117个样本)和验证集(6,014个样本),总数据量约15.6MB。每个数据样本由原始文本(source_text)、语言标识(language)和隐私掩码标注(privacy_mask)组成,其中privacy_mask标注了文本中隐私敏感信息的起始位置、结束位置、标签和原始值,适用于自然语言处理任务中的隐私保护或信息抽取研究。
This dataset is a text dataset designed for privacy-sensitive information annotation. It includes a training set with 54,117 samples and a validation set with 6,014 samples, with a total size of approximately 15.6 MB. Each data sample comprises the original text (source_text), language identifier (language), and privacy mask annotation (privacy_mask). The privacy_mask annotates the start position, end position, label and original value of the privacy-sensitive information in the text, and is applicable to research on privacy protection or information extraction in natural language processing tasks.




