遇见数据集

Spanish Morality Corpus

收藏
Zenodo2025-05-09 更新2026-05-26 收录
官方服务:

资源简介:

Description of the files included in this dataset This dataset contains the first publishable version of a corpus of Spanish-language online comments annotated for moral foundations, developed as part of the AMOR project. The annotations were carried out by trained participants using the Qualtrics platform and coordinated via Prolific. The comments were extracted from Spanish-speaking Reddit communities, filtered and manually curated to ensure linguistic quality, moral relevance, and geographical diversity. The dataset comprises two main components: Corpus Files (.json and .jsonl formats): These files include the annotated texts, each accompanied by metadata such as subreddit, author, and comment thread information. Annotations capture moral foundations (e.g., care, loyalty, authority) and include additional information on polarity (virtue or vice) and the annotator’s confidence level (low, medium, high). Three versions are provided for each annotation task: including all confidence levels, high confidence only, and medium+high confidence. Annotator Profiles (.json and .jsonl formats): These files describe the demographic and moral profile of each annotator using the MFQ30 (Moral Foundations Questionnaire) adapted to Spanish. Identifiers are anonymized for privacy. The dataset is designed for use in computational linguistics and social science research on moral language and value expression in Spanish-speaking online discourse.

提供机构:
Zenodo
创建时间:
2025-05-05
二维码
社区交流群
二维码
科研交流群
商业服务