基于r/Stalking子论坛的受害者提问数据集
收藏资源简介:
该数据集由德克萨斯大学阿灵顿分校与东北大学的研究团队构建,旨在为技术辅助滥用(TFA)受害者支持系统的评估提供真实、以受害者为导向的查询数据。数据集内容源于知名在线社区r/Stalking长达十年(2014-2024年)的公开讨论,通过定性编码与监督分类方法,从受害者自述中提取了2797条涉及11类技术滥用(如监控追踪、未经授权访问、社交媒体滥用等)的真实求助问题。其创建过程融合了人工标注与大语言模型辅助分类,确保了数据在反映受害者动态信息需求方面的真实性与代表性。该数据集主要应用于评估在线支持系统(如网络搜索、同伴支持论坛和对话式AI)在回应TFA受害者时的指导质量与安全性,旨在揭示现有数字支持生态系统的关键缺陷,并推动以安全为中心的未来支持技术设计与部署。
This dataset was developed by a research team from The University of Texas at Arlington and Northeastern University, with the objective of providing authentic, victim-centric query data for evaluating support systems for Technical Facilitated Abuse (TFA) victims. The dataset is sourced from a decade-long (2014–2024) public discussion on the popular Reddit subreddit r/Stalking. Using qualitative coding and supervised classification methods, 2,797 real help-seeking queries covering 11 categories of technical abuse—including monitoring and tracking, unauthorized access, social media abuse, and others—were extracted from victims’ self-reported accounts. The dataset’s creation process combines manual annotation and Large Language Model (LLM)-assisted classification, ensuring the data’s authenticity and representativeness in reflecting victims’ dynamic information needs. This dataset is primarily utilized to assess the guidance quality and safety of online support systems, such as web search tools, peer support forums, and conversational AI, when responding to TFA victims. It aims to uncover critical flaws in existing digital support ecosystems and promote the design and deployment of safety-focused future support technologies.




