遇见数据集

MAPPING THE DIGITAL SCIENTIFIC DEBATE ON AI: DISCIPLINARY NARRATIVES, PLATFORM DYNAMICS, AND THE ROLE OF MEDIA AND COMMUNICATION

收藏
Zenodo2026-04-08 更新2026-05-26 收录
官方服务:

资源简介:

This dataset accompanies the article Mapping the Digital Scientific Debate on AI: Disciplinary Narratives, Platform Dynamics, and the Role of Media and Communication. It contains the final analytical sample of 6,215 social media posts discussing artificial intelligence in relation to scientific research, drawn from an initial corpus of 9,844 posts published during the first half of 2025 across Instagram, X, TikTok, LinkedIn, and Bluesky. The dataset was designed to support transparent and reusable research on how AI is publicly debated as a scientific tool across disciplines and platforms. To maximize privacy protection, it does not include raw post text, usernames, profile data, URLs, or other direct identifiers. Instead, it only contains LLM-inferred and derived analytical fields, following a strict principle of irreversible anonymization and data minimization. This makes the dataset suitable for secondary analysis of discursive patterns while substantially reducing re-identification risks. The annotation workflow combined computational content analysis with Large Language Models. First, posts were classified into OECD/FORD research fields, with posts not clearly related to academic research excluded from the final released dataset. Second, the valid research-related posts were coded through a closed codebook for dimensions such as disciplinary focus, content type, general topic, AI sentiment, AI stance, risks, opportunities, audience, framing, and synthetic discourse indicators. The final dataset therefore captures structured interpretive variables rather than original social media content. This resource is intended for researchers interested in science communication, platform studies, AI discourse, computational social science, and the public understanding of science. Because the released file only contains inferred variables, it is especially useful for reproducible quantitative analyses of narrative patterns, disciplinary differences, framing strategies, and platform-specific dynamics without redistributing identifiable platform content. FORD field legend 0 = None / Not clearly research 1 = Natural sciences 2 = Engineering and technology 3 = Medical and health sciences 4 = Agricultural and veterinary sciences 5 = Social sciences 6 = Humanities and the arts

本数据集配套论文《绘制人工智能的数字化科学辩论图谱:学科叙事、平台动态及媒介传播的角色》。本数据集包含最终分析样本,共计6215条围绕与科学研究相关的人工智能展开讨论的社交媒体帖文,其原始语料库源自2025年上半年发布于Instagram、X、TikTok、LinkedIn及Bluesky平台的9844条初始帖文。 本数据集旨在支撑透明化、可复用的研究,以探究人工智能作为跨学科、跨平台的科学工具是如何被公众讨论的。为最大限度保护隐私,数据集未包含原始帖文文本、用户名、个人主页数据、URL或其他直接身份标识。取而代之的是,其仅包含经大语言模型(Large Language Model,LLM)推导生成的衍生分析字段,严格遵循不可逆匿名化与数据最小化原则。这使得该数据集既适用于话语模式的二次分析,又能大幅降低重识别风险。 该标注工作流将计算式内容分析与大语言模型相结合。首先,帖文被归类至经合组织(OECD)/FORD研究领域,与学术研究关联不明确的帖文将被排除在最终发布的数据集中。其次,针对有效的科研相关帖文,研究人员通过封闭编码手册对多维度内容进行编码,涵盖学科聚焦、内容类型、核心主题、人工智能情感倾向、人工智能立场、风险、机遇、受众、框架及合成话语指标等维度。因此,最终数据集所承载的是结构化的解释性变量,而非原始社交媒体内容。 本资源面向关注科学传播、平台研究、人工智能话语、计算社会科学以及公众对科学的理解的研究人员。由于发布文件仅包含推导生成的变量,其尤其适用于对叙事模式、学科差异、框架策略及平台特定动态开展可复现的定量分析,且无需重新分发可识别的平台内容。 FORD字段图例 0 = 无/与研究关联不明确 1 = 自然科学 2 = 工程与技术 3 = 医学与健康科学 4 = 农业与兽医学 5 = 社会科学 6 = 人文与艺术

提供机构:
Zenodo
创建时间:
2026-04-08
二维码
社区交流群
二维码
科研交流群
商业服务