遇见数据集

Dataset for "Enhancing multi-label emotion analysis in Indonesian social media with emoji-aware text representations"

收藏
Zenodo2026-01-12 更新2026-05-26 收录
官方服务:

资源简介:

Creator Amalia Amalia1*, Maya Silvi Lydia1, Rahmi Putri Rangkuti2, Farhan Purwanto Marulitua3, Fikri Hanif3, Mardanan Fitra3 , Dani Gunawan4 1 Department of Computer Science, Universitas Sumatera Utara, Indonesia 2 Department of Psychology, Universitas Sumatera Utara, Indonesia 3 Department of Data Science and Artificial Intelligence, Universitas Sumatera Utara, Indonesia 4 Department of Information Technology, Universitas Sumatera Utara, Indonesia Funding This research was supported by the Directorate of Research, Technology, and Community Service (DRTPM), Ministry of Education, Culture, Research, and Technology of the Republic of Indonesia, under the Fundamental Research Grant Scheme (Regular) 2025, based on Decree No. 0419/C3/DT.05.00/2025 and Contract/Agreement No. 112/C3/DT.05.00/PL/2025. Description The dataset was constructed from Indonesian social media content, integrating both textual data and paralinguistic signals in the form of emojis. In total, it consists of 139,414 instances annotated for Text Emotion Analysis (TEA) based on Plutchik’s emotion model. Unlike single-label corpora, this dataset supports multi-label classification, where a single post may express more than one emotion simultaneously. For example, text accompanied by multiple emojis (e.g., 😍 and 😢) can convey both joy and sadness, resulting in overlapping emotion categories. To enrich the emotional representation, a vocabulary of 1,040 unique emojis was incorporated, capturing supportive, contrastive, or even sarcastic emotional cues. This makes the dataset a valuable resource for exploring multimodal and multi-label emotion analysis in Indonesian, a low-resource language where high-quality annotated datasets are scarce. The dataset is specifically designed to benchmark models that integrate paralinguistic signals and to evaluate the robustness of LLM-based TEA systems. This dataset was developed for and introduced in the study titled Emoji-Aware Multimodal Representations for Multi-Label Emotion Analysis in Indonesian Social Media. Annotation Methodology The annotation process followed a multi-stage pipeline designed to balance scale and precision. Initial labels were assigned using a rule-based approach based on a Plutchik seed-word lexicon. To capture contextual nuances and linguistic subtleties, these labels were subsequently refined using a Large Language Model (LLM) with specialized prompts. Finally, the automated labels underwent human validation and correction by three expert annotators to ensure the highest degree of accuracy. Label Definitions and Annotation Guidelines To ensure consistency across 139,414 instances, annotators adhered to a strict set of definitions and decision rules tailored for Indonesian social media discourse. The dataset employs the 8 standard Plutchik emotions, adapted for Indonesian social media context. Joy: Expressions of happiness, laughter (e.g., "wkwk"), or positive appreciation. Trust: Expressions of faith, agreement, or admiration towards a figure/idea. Fear: Expressions of anxiety, worry, or feeling threatened. Surprise: Expressions of shock or unexpected realization (e.g., "kaget", "waduh"). Sadness: Expressions of grief, disappointment, or loss. Disgust: Expressions of revulsion, strong dislike, or feeling offended. Anger: Expressions of rage, frustration, or hostility. Anticipation: Expressions of hope, waiting, or planning for the future. Annotation Rules Annotations were performed by jointly analyzing textual content and emojis as a single semantic unit. Emojis were treated as meaningful signals that could clarify ambiguity, intensify emotional expression, or alter the literal interpretation of the text. In cases where textual and emoji cues conveyed contrasting tones, such as negative statements paired with positive or playful emojis, the annotation prioritized the intended meaning inferred from context rather than a surface-level reading. For example, a complaint accompanied by a laughter emoji (e.g., 🤣) could reflect mockery and therefore be annotated with multiple emotions, such as Disgust and Joy, when both were clearly implied. Although the annotation scheme allowed multi-label assignments, annotators were instructed to select only emotions that were explicitly evident in the post. Secondary emotions were included only when they played a substantial role in the overall expression; for instance, Anger was added to a Disgust-dominant post only if clear aggressive or hostile language was present. All annotations were conducted strictly from the author’s perspective, aiming to capture the emotions expressed by the writer rather than the emotions potentially elicited in the reader. Quality Control and Inter-Annotator Agreement To verify the consistency of the annotation logic, a pilot study was conducted on a stratified sample of 500 instances. All three annotators independently labeled this subset, and the inter-annotator agreement was measured using Fleiss’ Kappa. The results indicated substantial agreement, with an average Fleiss' Kappa of 0.6027. The specific scores per emotion are as follows: Emotion Category Fleiss' Kappa Fear 0.6738 Disgust 0.6518 Sadness 0.6201 Surprise 0.6102 Trust 0.6047 Anger 0.5692 Joy 0.5631 Anticipation 0.5285 Following the pilot study, the remaining instances were distributed equally among the three annotators for validation. For instances identified as ambiguous, a consensus-based approach was utilized where annotators discussed the context to reach a final agreement. Data Partitioning For benchmarking and model training, the dataset was partitioned using an Iterative Train-Test Split. This technique was specifically chosen to preserve the label distribution across both the training and testing sets, accounting for the multi-label characteristics of the Indonesian social media data.

提供机构:
Zenodo
创建时间:
2026-01-08
二维码
社区交流群
二维码
科研交流群
商业服务