IMPACTeen
收藏资源简介:
IMPACTeen是由弗罗茨瓦夫理工大学创建的青少年社会影响力文本数据集,专注于描绘青少年在人际、媒体和数字环境中的社会影响场景。该数据集包含1,021个文本,涵盖对话和叙述形式,总计5,100条人工标注记录,数据源通过受控的大语言模型生成并结合人工两阶段编辑验证以确保真实性。其构建过程基于SITT分类法选取20种社会影响技术,通过维度约束生成上下文向量并人工优化。该数据集旨在支持社会影响力检测、标注者分歧分析、跨语言建模以及大语言模型的训练与评估,为解决青少年群体社会影响机理的跨学科研究提供关键资源。
IMPACTeen is a teenage social impact text dataset developed by Wrocław University of Science and Technology, focusing on depicting social impact scenarios involving adolescents in interpersonal, media, and digital environments. This dataset includes 1,021 texts in both conversational and narrative formats, totaling 5,100 manually annotated records. Its data sources were generated via controlled large language models (LLMs) and validated through two-stage manual editing to ensure authenticity. The construction pipeline selects 20 social influence techniques based on the SITT taxonomy, generates context vectors via dimensional constraints, and optimizes them manually. This dataset aims to support social impact detection, annotator disagreement analysis, cross-lingual modeling, as well as the training and evaluation of large language models (LLMs), providing a critical resource for interdisciplinary research investigating the mechanisms of social influence among adolescents.
数据集概述:IMPACTeen
英文全称:IMPACTeen: Intentions, Manipulation, Persuasion, Annotations, and Consequences in Teen Communication Dataset
核心主题:这是一个聚焦于青少年语境下文本社交影响场景的数据集,涵盖意图、操纵、说服、标注及后果。
数据规模与构成:
- 包含 1,021 段文本。
- 包含 5,100 条个人标注记录。
- 为社交影响技术提供 黄金标签。
标注视角:每段文本均从五个不同视角进行标注:
- 青少年
- 父母
- 心理学家
- 传播学专家
- 教师
构建方法:
- 通过受约束的大型语言模型(LLM)生成。
- 随后经过两步人工编辑和验证阶段,以确保青少年语境下的真实性。
标注维度:多维度标注覆盖了以下方面:
- 影响力存在与否
- 影响技术
- 意图
- 后果
- 抵抗性
- 反应
- 标注置信度
适用研究方向:
- 社交影响检测
- 标注者分歧研究
- 跨语言建模
- 语言模型的训练与评估
语言:原始数据集以波兰语创建,并附带对应的英文版本。
所属学科:计算机科学 > 计算与语言(cs.CL);人工智能(cs.AI)




