遇见数据集

Human–AI Conflict Discourse Dataset

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

Human–AI Conflict and Constraint Discourse Dataset (YouTube, 2010–2024) is a large-scale corpus of 209,228 English-language YouTube comments related to artificial intelligence, collected via the YouTube Data API from videos identified using the term “Artificial Intelligence.” The dataset includes cleaned textual content, publication metadata, sentiment classifications generated using a pretrained DistilBERT model, and topic cluster assignments derived from Latent Dirichlet Allocation, mapping public discourse to key AI constraint mechanisms (parametric reductionism, agency transference, and regulated expression). The corpus captures naturalistic public perceptions of AI across personal and institutional contexts and is intended to support research on human–AI interaction, AI ethics, and computational analysis of digital discourse.

人机冲突与约束话语数据集(YouTube平台,2010-2024年)为规模庞大的语料库,包含209228条与人工智能相关的英文YouTube评论,通过YouTube数据API(YouTube Data API)从以"Artificial Intelligence"为检索词识别出的视频中采集所得。本数据集涵盖经清洗的文本内容、发布元数据、基于预训练DistilBERT模型生成的情感分类结果,以及由潜在狄利克雷分配(Latent Dirichlet Allocation)推导得到的主题聚类分配结果,将公共话语映射至核心人工智能约束机制:参数化还原论、主体性转移与规范化表达。该语料库囊括个人与制度场景下公众对人工智能的自然主义认知,旨在为人机交互、人工智能伦理学以及数字话语计算分析相关研究提供支撑。

创建时间:
2025-12-25
二维码
社区交流群
二维码
科研交流群
商业服务