遇见数据集

Dataset for "Enhancing multi-label emotion analysis in Indonesian social media with emoji-aware text representations"

收藏
Zenodo2026-01-12 更新2026-05-26 收录
官方服务:

资源简介:

Creator Amalia Amalia1*, Maya Silvi Lydia1, Rahmi Putri Rangkuti2, Farhan Purwanto Marulitua3, Fikri Hanif3, Mardanan Fitra3 , Dani Gunawan4 1 Department of Computer Science, Universitas Sumatera Utara, Indonesia 2 Department of Psychology, Universitas Sumatera Utara, Indonesia 3 Department of Data Science and Artificial Intelligence, Universitas Sumatera Utara, Indonesia 4 Department of Information Technology, Universitas Sumatera Utara, Indonesia Funding This research was supported by the Directorate of Research, Technology, and Community Service (DRTPM), Ministry of Education, Culture, Research, and Technology of the Republic of Indonesia, under the Fundamental Research Grant Scheme (Regular) 2025, based on Decree No. 0419/C3/DT.05.00/2025 and Contract/Agreement No. 112/C3/DT.05.00/PL/2025. Description The dataset was constructed from Indonesian social media content, integrating both textual data and paralinguistic signals in the form of emojis. In total, it consists of 139,414 instances annotated for Text Emotion Analysis (TEA) based on Plutchik’s emotion model. Unlike single-label corpora, this dataset supports multi-label classification, where a single post may express more than one emotion simultaneously. For example, text accompanied by multiple emojis (e.g., 😍 and 😢) can convey both joy and sadness, resulting in overlapping emotion categories. To enrich the emotional representation, a vocabulary of 1,040 unique emojis was incorporated, capturing supportive, contrastive, or even sarcastic emotional cues. This makes the dataset a valuable resource for exploring multimodal and multi-label emotion analysis in Indonesian, a low-resource language where high-quality annotated datasets are scarce. The dataset is specifically designed to benchmark models that integrate paralinguistic signals and to evaluate the robustness of LLM-based TEA systems. This dataset was developed for and introduced in the study titled Emoji-Aware Multimodal Representations for Multi-Label Emotion Analysis in Indonesian Social Media. Annotation Methodology The annotation process followed a multi-stage pipeline designed to balance scale and precision. Initial labels were assigned using a rule-based approach based on a Plutchik seed-word lexicon. To capture contextual nuances and linguistic subtleties, these labels were subsequently refined using a Large Language Model (LLM) with specialized prompts. Finally, the automated labels underwent human validation and correction by three expert annotators to ensure the highest degree of accuracy. Label Definitions and Annotation Guidelines To ensure consistency across 139,414 instances, annotators adhered to a strict set of definitions and decision rules tailored for Indonesian social media discourse. The dataset employs the 8 standard Plutchik emotions, adapted for Indonesian social media context. Joy: Expressions of happiness, laughter (e.g., "wkwk"), or positive appreciation. Trust: Expressions of faith, agreement, or admiration towards a figure/idea. Fear: Expressions of anxiety, worry, or feeling threatened. Surprise: Expressions of shock or unexpected realization (e.g., "kaget", "waduh"). Sadness: Expressions of grief, disappointment, or loss. Disgust: Expressions of revulsion, strong dislike, or feeling offended. Anger: Expressions of rage, frustration, or hostility. Anticipation: Expressions of hope, waiting, or planning for the future. Annotation Rules Annotations were performed by jointly analyzing textual content and emojis as a single semantic unit. Emojis were treated as meaningful signals that could clarify ambiguity, intensify emotional expression, or alter the literal interpretation of the text. In cases where textual and emoji cues conveyed contrasting tones, such as negative statements paired with positive or playful emojis, the annotation prioritized the intended meaning inferred from context rather than a surface-level reading. For example, a complaint accompanied by a laughter emoji (e.g., 🤣) could reflect mockery and therefore be annotated with multiple emotions, such as Disgust and Joy, when both were clearly implied. Although the annotation scheme allowed multi-label assignments, annotators were instructed to select only emotions that were explicitly evident in the post. Secondary emotions were included only when they played a substantial role in the overall expression; for instance, Anger was added to a Disgust-dominant post only if clear aggressive or hostile language was present. All annotations were conducted strictly from the author’s perspective, aiming to capture the emotions expressed by the writer rather than the emotions potentially elicited in the reader. Quality Control and Inter-Annotator Agreement To verify the consistency of the annotation logic, a pilot study was conducted on a stratified sample of 500 instances. All three annotators independently labeled this subset, and the inter-annotator agreement was measured using Fleiss’ Kappa. The results indicated substantial agreement, with an average Fleiss' Kappa of 0.6027. The specific scores per emotion are as follows: Emotion Category Fleiss' Kappa Fear 0.6738 Disgust 0.6518 Sadness 0.6201 Surprise 0.6102 Trust 0.6047 Anger 0.5692 Joy 0.5631 Anticipation 0.5285 Following the pilot study, the remaining instances were distributed equally among the three annotators for validation. For instances identified as ambiguous, a consensus-based approach was utilized where annotators discussed the context to reach a final agreement. Data Partitioning For benchmarking and model training, the dataset was partitioned using an Iterative Train-Test Split. This technique was specifically chosen to preserve the label distribution across both the training and testing sets, accounting for the multi-label characteristics of the Indonesian social media data.

数据集创建者 Amalia Amalia1*, Maya Silvi Lydia1, Rahmi Putri Rangkuti2, Farhan Purwanto Marulitua3, Fikri Hanif3, Mardanan Fitra3, Dani Gunawan4 1 印度尼西亚北苏门答腊大学计算机科学系 2 印度尼西亚北苏门答腊大学心理学系 3 印度尼西亚北苏门答腊大学数据科学与人工智能系 4 印度尼西亚北苏门答腊大学信息技术系 资助 本研究获得印度尼西亚共和国教育、文化、研究与技术部研究、技术与社区服务总局(DRTPM)资助,依托2025年度基础研究资助计划(常规类),依据第0419/C3/DT.05.00/2025号法令及第112/C3/DT.05.00/PL/2025号合同/协议执行。 数据集概述 本数据集源自印度尼西亚社交媒体内容,整合了文本数据与以表情符号为载体的副语言信号。数据集共包含139,414条标注样本,用于基于普卢特奇克(Plutchik)情绪模型的文本情绪分析(Text Emotion Analysis, TEA)。与单标签语料库不同,本数据集支持多标签分类,单条帖子可同时表达多种情绪。例如,搭配多个表情符号(如😍和😢)的文本可同时传递喜悦与悲伤,从而形成重叠的情绪类别。为丰富情绪表征,数据集纳入了包含1,040个独特表情符号的词汇表,可捕捉支持性、对比性乃至讽刺性的情绪线索。这使得该数据集成为探索印度尼西亚语(高质量标注数据集稀缺的低资源语言)多模态多标签情绪分析的宝贵资源。本数据集专为集成副语言信号的模型基准测试,以及评估基于大语言模型(Large Language Model, LLM)的文本情绪分析系统的鲁棒性而设计。 本数据集为发表于题为《面向印度尼西亚社交媒体多标签情绪分析的表情感知多模态表征》的研究而开发并推出。 标注方法论 标注流程遵循多阶段流水线设计,以兼顾标注规模与精度。初始标签采用基于普卢特奇克种子词词典的规则方法进行分配。为捕捉上下文细微差别与语言细节,后续使用带有专用提示词的大语言模型(LLM)对这些标签进行优化。最终,自动生成的标签经三位专家标注者进行人工验证与修正,以确保最高标注精度。 标签定义与标注指南 为确保139,414条样本的标注一致性,标注人员遵循针对印度尼西亚社交媒体话语定制的严格定义与决策规则。本数据集采用适配印度尼西亚社交媒体语境的8种标准普卢特奇克情绪类别: - 喜悦(Joy):表达幸福、欢笑(如“wkwk”)或正向赞赏 - 信任(Trust):表达对某一人物/观点的信念、认同或钦佩 - 恐惧(Fear):表达焦虑、担忧或受威胁的感受 - 惊讶(Surprise):表达震惊或意外的领悟(如“kaget”“waduh”) - 悲伤(Sadness):表达悲痛、失望或失落 - 厌恶(Disgust):表达反感、强烈厌恶或被冒犯的感受 - 愤怒(Anger):表达暴怒、挫败或敌意 - 期待(Anticipation):表达希望、等待或对未来的规划 标注规则 标注需将文本内容与表情符号作为单一语义单元联合分析。表情符号被视为可澄清歧义、强化情绪表达或改变文本字面含义的有效信号。当文本与表情符号传递的语气存在冲突时(如负面陈述搭配正向或戏谑表情),标注优先考虑从上下文推断的作者意图,而非表面解读。例如,搭配笑到流泪表情(🤣)的抱怨可能带有嘲讽意味,此时若同时隐含厌恶与喜悦情绪,则可标注为多标签类别。尽管标注方案支持多标签分配,但标注人员仅需选择帖子中明确显现的情绪。次级情绪仅当其在整体表达中发挥实质性作用时才可纳入;例如,仅当以厌恶为主的帖子中存在明确的攻击性或敌意语言时,才可添加愤怒标签。所有标注严格从作者视角出发,旨在捕捉作者表达的情绪,而非读者可能产生的情绪。 质量控制与标注者间一致性 为验证标注逻辑的一致性,针对分层抽样的500条样本开展预实验。三位标注者独立对该子集进行标注,采用弗莱斯Kappa系数(Fleiss’ Kappa)衡量标注者间一致性。结果显示一致性良好,平均弗莱斯Kappa值为0.6027。各情绪类别具体得分如下: | 情绪类别 | 弗莱斯Kappa系数 | | --- | --- | | 恐惧 | 0.6738 | | 厌恶 | 0.6518 | | 悲伤 | 0.6201 | | 惊讶 | 0.6102 | | 信任 | 0.6047 | | 愤怒 | 0.5692 | | 喜悦 | 0.5631 | | 期待 | 0.5285 | 预实验结束后,剩余样本均分给三位标注者进行验证。对于存在歧义的样本,采用基于共识的方法,由标注者共同讨论上下文以达成最终一致意见。 数据划分 为开展基准测试与模型训练,本数据集采用迭代训练测试划分(Iterative Train-Test Split)方法进行划分。该方法经专门选择,可保留训练集与测试集的标签分布,适配印度尼西亚社交媒体数据的多标签特性。

提供机构:
Zenodo
创建时间:
2025-10-08
二维码
社区交流群
二维码
科研交流群
商业服务