遇见数据集

Snapshots of a Culture War: Dataset of X/Twitter and Reddit Posts from Conservative Christian and LGBTQIA+ Youth Issues Discourse Across Four Time Periods, with Sentiment Scores

收藏
Mendeley Data2026-09-08 收录
官方服务:

资源简介:

This dataset consists of 1159 entries scraped from Reddit and X/Twitter between February 10 and March 20, 2026, including at least 100 posts from each platform representing each of four different political/cultural/historical/technological moments in the online cultural struggle between the interests of LGBTQIA+ youth and those of conservative Christian ideologies and institutions. Posts were collected from before the Trump era, during the Trump Era but before the COVID-19 lockdowns, during the COVID Era, and after the COVID Era. Each post has been assigned three sentiment codes: One by a machine learning approach created for general language, one by a machine learning approach created specifically for X/Twitter, and one by a social scientist who researches LGBTQIA+ youths’ issues with conservative Christianity, who has previously published research that involved coding short-format text data. To keep work on this dataset within the realm of activity that is not human subjects research, the dataset does not include the body text of posts marked as deleted or removed in their respective databases. For the sake of technical completeness, the small percentage of posts found by our algorithms that were not really part of the discourse (e.g., product advertisements, “AI slop”) were left in, along with the sentiment categories our machine learning approaches assigned to them. This dataset may be useful to investigators of LGBTQIA+ youths’ real-lived experiences with online discourse involving conservative Christian ideologies and institutions, and to computer scientists interested in training sentiment analysis models to address politically charged issues.

本数据集包含1159条采集自2026年2月10日至3月20日期间Reddit与X/Twitter平台的帖文数据,两个平台各至少提供100条帖文,对应围绕性少数群体(LGBTQIA+)青年权益与保守基督教意识形态及机构利益展开的在线文化斗争中的四类不同政治、文化、历史与技术节点。帖文采集覆盖四个阶段:特朗普执政前、特朗普执政时期但新冠疫情封控前、新冠疫情时期以及新冠疫情后。每条帖文均被赋予三类情感编码:第一类由面向通用语言开发的机器学习方法生成,第二类由专为X/Twitter平台开发的机器学习方法生成,第三类由一名研究性少数群体(LGBTQIA+)青年与保守基督教相关议题的社会科学家标注——该学者此前已发表过涉及短格式文本数据编码的研究成果。为使本数据集的相关研究活动免于落入人类受试者研究范畴,数据集未收录各平台数据库中被标记为删除或移除的帖文正文。出于技术完整性的考量,算法抓取到的少量不属于该讨论范畴的帖文(如商业广告、‘AI垃圾内容’)被保留,同时保留了机器学习方法为其分配的情感分类标签。本数据集可服务于两类研究者:一类是探究性少数群体(LGBTQIA+)青年在涉及保守基督教意识形态及机构的在线话语中的真实生存体验的研究者,另一类是致力于训练针对政治敏感议题的情感分析模型的计算机科学研究者。

创建时间:
2026-08-17
二维码
社区交流群
二维码
科研交流群
商业服务