Spanish Morality Corpus
收藏资源简介:
Description of the files included in this dataset This dataset contains the first publishable version of a corpus of Spanish-language online comments annotated for moral foundations, developed as part of the AMOR project. The annotations were carried out by trained participants using the Qualtrics platform and coordinated via Prolific. The comments were extracted from Spanish-speaking Reddit communities, filtered and manually curated to ensure linguistic quality, moral relevance, and geographical diversity. The dataset comprises two main components: Corpus Files (.json and .jsonl formats): These files include the annotated texts, each accompanied by metadata such as subreddit, author, and comment thread information. Annotations capture moral foundations (e.g., care, loyalty, authority) and include additional information on polarity (virtue or vice) and the annotator’s confidence level (low, medium, high). Three versions are provided for each annotation task: including all confidence levels, high confidence only, and medium+high confidence. Annotator Profiles (.json and .jsonl formats): These files describe the demographic and moral profile of each annotator using the MFQ30 (Moral Foundations Questionnaire) adapted to Spanish. Identifiers are anonymized for privacy. The dataset is designed for use in computational linguistics and social science research on moral language and value expression in Spanish-speaking online discourse. AMOR-Corpus V2 – Format and Annotator Profiling Description The AMOR-Corpus V2 is a curated, high-quality dataset of Spanish-language Reddit comments annotated for moral content. It was created as part of the AMOR project, focused on affective and moral reasoning in online discourse. The dataset is composed of two structured JSON-based files: 1. AMOR-Corpus_V2-high.jsonl – Corpus and Annotations This file contains Reddit comments annotated for the presence of moral foundations, based on an expanded version of Moral Foundations Theory (MFT). Each line in this JSON Lines file represents: - id: A unique identifier for the comment.- text: The original Reddit comment in Spanish.- annotations: A list of five independent annotations per comment, each including: - The moral foundation(s) perceived (if any). - Subtypes (Virtue/Vice distinctions). - Annotator confidence level (Low, Medium, High). Annotations were collected through the Qualtrics platform, and annotators were recruited via Prolific. Each task included 60 comments and was annotated by 5 different individuals. Annotators had the option to mark “None” if no moral content was identified. Manual pre-filtering of texts ensured content was:- Written in Spanish,- Interpretable without deep context,- Potentially moral in nature,- Free of extreme offensive language. 2. AMOR-Corpus_V2-annotators.json – Annotator Metadata This file contains anonymized demographic and psychometric profiles of all annotators. Each entry includes:- ID: A unique anonymized user ID.- genre, age, political orientation, income, religious: Self-reported demographic information.- Responses to the Moral Foundations Questionnaire (MFQ), described below. Moral Foundations Questionnaire (MFQ) To assess how personal values may influence moral annotation behavior, each annotator completed the Moral Foundations Questionnaire (MFQ), a validated psychometric instrument based on MFT. The MFQ measures sensitivity across five foundational moral domains: 1. Care/Harm2. Fairness/Cheating3. Loyalty/Betrayal4. Authority/Subversion5. Purity/Degradation Each foundation includes:- Relevance items (e.g., “Whether or not someone suffered emotionally”), measuring importance,- Judgment items (e.g., “Compassion for those who are suffering is the most crucial virtue”), measuring agreement. Annotators responded using a 6-point Likert scale. Item codes in the data are prefixed as follows:- CA_# = Care- EQ_# = Fairness (EQ = Equality)- LO_# = Loyalty- AU_# = Authority- PU_# = Purity- PR_# = Proportionality-related fairness The inclusion of MFQ scores enables researchers to analyze how individual moral orientations influence perception and annotation of moral language. More information on MFQ can be found at:- MFQ1 (original): https://moralfoundations.org/questionnaires/- MFQ2 (updated): https://yourmorals.org/
本数据集所含文件说明 本数据集包含首个可公开出版的西班牙语在线评论语料库版本,该语料库针对道德基础进行标注,作为AMOR项目的一部分开发完成。标注工作由经过培训的参与者通过Qualtrics平台完成,并通过Prolific进行协调。评论数据取自西班牙语区Reddit社区,经过筛选与人工审核整理,以确保语言质量、道德相关性与地域多样性。 本数据集包含两大核心组成部分: ### 语料文件(.json与.jsonl格式) 此类文件包含已标注的文本,每条文本均附带元数据,如所属子社区、作者与评论线程信息。标注内容涵盖道德基础(如关爱、忠诚、权威等),同时包含极性(美德或恶行)标注以及标注者的置信度等级(低、中、高)。每个标注任务提供三个版本:涵盖所有置信度等级的版本、仅保留高置信度的版本,以及中高置信度组合的版本。 ### 标注者档案(.json与.jsonl格式) 此类文件通过适配西班牙语的MFQ30(道德基础问卷)描述每位标注者的人口统计学与道德画像。为保护隐私,所有标识符均已匿名化。 本数据集旨在服务于计算语言学与社会科学领域的研究,用于分析西班牙语在线话语中的道德语言与价值表达。 AMOR语料库V2——格式与标注者档案说明 AMOR-Corpus V2是一个经过精心整理的高质量西班牙语Reddit评论数据集,针对道德内容进行标注。该数据集作为AMOR项目的一部分创建,该项目聚焦于在线话语中的情感与道德推理。数据集由两个基于JSON的结构化文件组成: 1. AMOR-Corpus_V2-high.jsonl——语料库与标注集 该文件包含基于扩展版道德基础理论(Moral Foundations Theory, MFT)标注了道德基础存在性的Reddit评论。此JSON Lines文件的每一行代表: - id:评论的唯一标识符 - text:西班牙语原文Reddit评论 - annotations:每条评论的5条独立标注列表,每条标注包含: - 感知到的道德基础(若有) - 子类型(美德/恶行区分) - 标注者置信度等级(低、中、高) 标注工作通过Qualtrics平台收集,标注者通过Prolific招募。每个标注任务包含60条评论,由5名不同的标注者完成。标注者可选择标记“无”以表示未识别到道德内容。文本经过人工预筛选,确保: - 为西班牙语文本 - 无需深层上下文即可理解 - 具有潜在道德属性 - 无极端冒犯性语言 2. AMOR-Corpus_V2-annotators.json——标注者元数据 该文件包含所有标注者的匿名化人口统计学与心理测量学档案。每个条目包含: - ID:唯一的匿名化用户ID - genre(性别)、age(年龄)、political orientation(政治倾向)、income(收入)、religious(宗教信仰):自我报告的人口统计学信息 - 道德基础问卷(Moral Foundations Questionnaire, MFQ)的作答结果,详见下文。 道德基础问卷(MFQ) 为评估个人价值观如何影响道德标注行为,每位标注者均完成了基于MFT的经过验证的心理测量工具——道德基础问卷(MFQ)。MFQ可测量五个核心道德领域的敏感度: 1. 关爱/伤害 2. 公平/欺骗 3. 忠诚/背叛 4. 权威/颠覆 5. 纯洁/堕落 每个道德领域包含: - 相关性条目(例如“某人是否在情感上遭受痛苦”),用于测量对该领域的重视程度 - 判断条目(例如“对受苦者的同情是最重要的美德”),用于测量对该表述的认同度 标注者采用6点李克特量表进行作答。数据中的条目代码前缀规则如下: - CA_# = 关爱维度 - EQ_# = 公平维度(EQ = Equality) - LO_# = 忠诚维度 - AU_# = 权威维度 - PU_# = 纯洁维度 - PR_# = 与比例公平相关的条目 MFQ分数的纳入可使研究者分析个体道德取向如何影响道德语言的感知与标注。 更多关于MFQ的信息可通过以下链接获取: - 原版MFQ1:https://moralfoundations.org/questionnaires/ - 更新版MFQ2:https://yourmorals.org/



