Cognitive Dataset (认知数据集)
收藏资源简介:
一个大规模双语(中文和英文)数据集,用于研究新闻文章用户评论中表达的人格特质和认知偏差。该数据集包含新闻文章与人工撰写评论的配对,每个评论都标注了大五人格特质维度或认知偏差标签,支持文本人格检测、认知偏差识别与分类、观点挖掘和跨语言心理语言建模等研究。
A large-scale bilingual (Chinese and English) dataset designed for investigating personality traits and cognitive biases expressed in user comments on news articles. This dataset comprises paired samples of news articles and human-written comments, with each comment annotated with Big Five personality trait dimensions or cognitive bias labels. It supports research directions including textual personality detection, cognitive bias recognition and classification, opinion mining, and cross-linguistic psycholinguistic modeling.
数据集概述
认知数据集(Cognitive Dataset) 是一个大规模、中英双语的数据集,用于研究用户对新闻评论中表达的人格特质和认知偏差。
核心信息
- 目的:支持人格检测、认知偏差识别、意见挖掘、跨语言心理语言学建模、可解释AI等研究。
- 语言:中文和英文。
- 数据构成:包含新闻文章、对应的人类评论、每个评论的标签(大五人格维度或认知偏差)、原因解释和摘要。
数据集统计
| 子集 | 任务 | 标签数量 | 每种语言条目数 | 总条目数 |
|---|---|---|---|---|
| PS | 人格特质 | 10 | 11,400 | 22,800 |
| CB | 认知偏差 | 38 | 43,320 | 86,640 |
| News | 原始新闻文章 | — | 1,140 | 1,140 |
- 数据划分:80% 训练集 / 20% 测试集,每个标签类别保持均衡(每个标签1,140条)。
任务定义
PS — 人格特质(大五人格模型)
基于大五(OCEAN)人格模型,每个评论标注了特质维度及其效价(积极或消极):
| 编号 | 中文标签 | 英文标签 | 维度 | 效价 |
|---|---|---|---|---|
| 0 | 正向外倾性 | Positive Extraversion | 外倾性 | 积极 |
| 1 | 负向外倾性 | Negative Extraversion | 外倾性 | 消极 |
| 2 | 正向宜人性 | Positive Agreeableness | 宜人性 | 积极 |
| 3 | 负向宜人性 | Negative Agreeableness | 宜人性 | 消极 |
| 4 | 正向尽责性 | Positive Conscientiousness | 尽责性 | 积极 |
| 5 | 负向尽责性 | Negative Conscientiousness | 尽责性 | 消极 |
| 6 | 正向神经质 | Positive Neuroticism | 神经质 | 积极 |
| 7 | 负向神经质 | Negative Neuroticism | 神经质 | 消极 |
| 8 | 正向开放性 | Positive Openness | 开放性 | 积极 |
| 9 | 负向开放性 | Negative Openness | 开放性 | 消极 |
共10个类别,每类1,140条。“正向神经质”指情绪稳定(低神经质),“负向神经质”指高神经质。
CB — 认知偏差
评论被标注为38种认知偏差中的一种。每种偏差有1,140条:
| 编号 | 中文标签 | 英文标签 |
|---|---|---|
| 0 | 确认偏差 | Confirmation Bias |
| 1 | 锚定偏差 | Anchoring Bias |
| 2 | 框架效应 | Framing Effect |
| 3 | 选择性感知偏差 | Selective Perception Bias |
| 4 | 基本归因错误 | Fundamental Attribution Error |
| 5 | 真相幻觉效应 | Illusory Truth Effect |
| 6 | 信念偏差 | Belief Bias |
| 7 | 可获得性偏差 | Availability Bias |
| 8 | 损失厌恶 | Loss Aversion |
| 9 | 过度自信偏差 | Overconfidence Bias |
| 10 | 动机性推理 | Motivated Reasoning |
| 11 | 防御性归因 | Defensive Attribution |
| 12 | 群体内偏差 | In-group Bias |
| 13 | 现状偏差 | Status Quo Bias |
| 14 | 邓宁-克鲁格效应 | Dunning-Kruger Effect |
| 15 | 选择性曝光 | Selective Exposure |
| 16 | 认知失调 | Cognitive Dissonance |
| 17 | 道德运气 | Moral Luck |
| 18 | 敌意媒体效应 | Hostile Media Effect |
| 19 | 意识形态偏差 | Ideological Bias |
| 20 | 光环效应 | Halo Effect |
| 21 | 逆火效应 | Backfire Effect |
| 22 | 错误的共识 | False Consensus |
| 23 | 天真现实主义 | Naïve Realism |
| 24 | 外群体同质性偏差 | Out-group Homogeneity Bias |
| 25 | 显著性偏差 | Salience Bias |
| 26 | 购买后合理化 | Post-Purchase Rationalization |
| 27 | 向下比较偏差 | Downward Comparison Bias |
| 28 | 公正世界假说 | Just-World Hypothesis |
| 29 | 可用性级联 | Availability Cascade |
| 30 | 幸存者偏差 | Survivorship Bias |
| 31 | 记忆偏差 | Memory Bias |
| 32 | 防御性偏差 | Defensive Bias |
| 33 | 地域偏差 | Regional Bias |
| 34 | 朴素犬儒主义 | Naïve Cynicism |
| 35 | 第三人称效应 | Third-Person Effect |
| 36 | 顺序效应 | Order Effect |
| 37 | 成长信念 | Growth Mindset |
数据格式
每条记录为一个JSON对象,包含以下字段:
| 字段 | 类型 | 描述 |
|---|---|---|
title |
string | 新闻文章标题 |
news |
string | 新闻文章全文 |
comment |
string | 针对新闻的人类评论 |
reason |
string | 解释为什么该评论反映给定标签 |
summary |
string | 评论的简明摘要 |
label |
string | 标注的标签(人格特质或认知偏差) |
number |
integer | 标签类别内的索引号(0-1139) |
许可与引用
- 许可:仅供研究使用。
- 引用:请引用该仓库的GitHub地址:https://github.com/your-repo/cognitive-dataset




