forePLay
收藏资源简介:
forePLay是一个专门为波兰语情色内容检测设计的手动标注数据集,由NASK国家研究所创建。该数据集包含24,768条句子,涵盖了从在线小说和波兰文学作品中提取的内容,具有多维度的标注体系,包括模糊性、暴力和社会不可接受性等维度。数据集的创建过程包括从不同来源的文本中进行系统采样,并进行详细的预处理和标注,确保了数据集的多样性和代表性。该数据集主要用于开发语言感知的内容审核系统,旨在解决非英语情色内容检测的挑战,特别是在形态复杂的语言中。
forePLay is a manually annotated dataset specifically designed for Polish-language erotic content detection, created by NASK, the National Research Institute. Comprising 24,768 sentences, the dataset features content extracted from online fiction and Polish literary works, and employs a multi-dimensional annotation framework covering ambiguity, violence, social unacceptability, and other relevant dimensions. The dataset development process includes systematic sampling of texts from diverse sources, followed by rigorous preprocessing and annotation, which guarantees the dataset's diversity and representativeness. This dataset is primarily utilized for developing language-aware content moderation systems, with the goal of addressing the challenges in non-English erotic content detection, particularly in morphologically complex languages.




