BAED: A Comprehensive Bengali Emotion Dataset with Transformer-based Evaluation
收藏资源简介:
The Bengali Annotated Emotion Dataset (BAED) draws its content from Bengali novels which form the basis for its emotional expression and cultural element identification capabilities that other NLP datasets do not possess. The system divides into seven categories which enable complete multi-class affective state classification through anger, disgust, fear, joy, sadness, surprise and anticipation. Drawing from both dialogue and narration, BAED offers a rich basis for studying human emotions in text.The system has multiple uses in computational linguistics and sentiment and stylistic analysis and cultural and psychological research and dialogue-level emotion detection and benchmarking emotion classification models with potential future applications in newspaper and social media and conversational data. The dataset is organized into seven balanced classes, each representing a different emotion domain: Anger:500 data Disgust:500 data Fear:500 data Joy:500 data Sadness:500 data Surprise:500 data Anticipation:500 data Total Number of Data: 3,500 Language: Bangla File Format: CSV file
孟加拉语标注情感数据集(Bengali Annotated Emotion Dataset,以下简称BAED)的内容源自孟加拉语小说,其情感表达与文化元素识别能力为该数据集独有,亦是其他自然语言处理(Natural Language Processing,简称NLP)数据集所不具备的。该数据集的分类体系涵盖七大类别,可实现完整的多分类情感状态分类,覆盖愤怒、厌恶、恐惧、喜悦、悲伤、惊讶与期待七种情绪。该数据集同时取材于对话与叙事文本,为文本中的人类情感研究提供了丰富的支撑基础。该数据集可应用于计算语言学、情感与文体分析、文化与心理学研究、对话级情感检测以及情感分类模型的基准测试等多个场景,未来还可拓展应用于报纸文本、社交媒体数据与对话数据相关任务。 该数据集被划分为七个平衡类别,每个类别对应独立的情感领域: 愤怒:500条数据 厌恶:500条数据 恐惧:500条数据 喜悦:500条数据 悲伤:500条数据 惊讶:500条数据 期待:500条数据 数据总条数:3500 语言:孟加拉语 文件格式:CSV文件



