Dataset for User Sensitive Data on Social Media
收藏资源简介:
This dataset comprises textual records collected from social media platforms, curated to support both academic researchers and industry practitioners in advancing privacy-aware natural language processing. Each record is annotated to indicate whether it constitutes a breach of user privacy or not. The dataset is designed to facilitate the development, training, and evaluation of machine learning models for privacy violation detection, enabling applications in areas such as content moderation, compliance monitoring, and online safety. By providing high-quality labeled data, this resource contributes to building safer AI systems that respect user confidentiality and protect sensitive information in digital interactions.
本数据集包含从社交媒体平台采集的文本记录,旨在为学术研究者与行业从业者推进隐私感知自然语言处理(Privacy-Aware Natural Language Processing)相关研究提供支撑。每条记录均标注了是否涉及用户隐私泄露情况。本数据集专为隐私违规检测(Privacy Violation Detection)相关机器学习模型的开发、训练与评估打造,可支撑内容审核、合规监测、在线安全等领域的应用落地。通过提供高质量标注数据,本数据集助力构建更安全的人工智能系统,在数字交互中尊重用户隐私机密并保护敏感信息。




