Spanish Morality Corpus
收藏资源简介:
Description of the files included in this dataset This dataset contains the first publishable version of a corpus of Spanish-language online comments annotated for moral foundations, developed as part of the AMOR project. The annotations were carried out by trained participants using the Qualtrics platform and coordinated via Prolific. The comments were extracted from Spanish-speaking Reddit communities, filtered and manually curated to ensure linguistic quality, moral relevance, and geographical diversity. The dataset comprises two main components: Corpus Files (.json and .jsonl formats): These files include the annotated texts, each accompanied by metadata such as subreddit, author, and comment thread information. Annotations capture moral foundations (e.g., care, loyalty, authority) and include additional information on polarity (virtue or vice) and the annotator’s confidence level (low, medium, high). Three versions are provided for each annotation task: including all confidence levels, high confidence only, and medium+high confidence. Annotator Profiles (.json and .jsonl formats): These files describe the demographic and moral profile of each annotator using the MFQ30 (Moral Foundations Questionnaire) adapted to Spanish. Identifiers are anonymized for privacy. The dataset is designed for use in computational linguistics and social science research on moral language and value expression in Spanish-speaking online discourse.



