quynong/rosenxt-v1-26-5
收藏资源简介:
该数据集包含多语言文本数据,每个样本具有语言标识、原始文本内容以及隐私掩码信息。隐私掩码用于标注文本中的隐私敏感信息,包括掩码的起始位置、结束位置、标签类型、固定状态和实际值。数据集采用五折交叉验证结构,共划分为5个训练集和5个测试集,每个训练集约8969个样本,每个测试集约2242个样本,总样本量约56000余条。适用于隐私保护、文本匿名化或信息抽取等自然语言处理任务。
This dataset contains multilingual text data, where each sample includes language identification, source text content, and privacy mask information. The privacy mask is used to annotate privacy-sensitive information in the text, covering mask start position, end position, label type, fix status, and actual value. The dataset is structured with five-fold cross-validation, comprising 5 training sets and 5 test sets. Each training set has approximately 8969 samples, and each test set has about 2242 samples, totaling over 56,000 samples. It is suitable for natural language processing tasks such as privacy protection, text anonymization, or information extraction.



