Emo Pillars
收藏资源简介:
Emo Pillars数据集由Pompeu Fabra University和Barcelona Supercomputing Center创建,是一个包含28种情感类别的情感分类数据集。该数据集通过利用大型语言模型Mistral-7b生成上下文丰富和上下文匮乏的句子,以支持细粒度的情感分类。数据集分为上下文丰富和上下文匮乏两部分,共计100K和300K条示例。该数据集用于微调预训练的编码器模型,以提高其在不同任务上的表现。
The Emo Pillars dataset, created by Pompeu Fabra University and Barcelona Supercomputing Center, is an emotion classification dataset encompassing 28 distinct emotion categories. To support fine-grained emotion classification, this dataset generates sentences with both context-rich and context-scarce content using the large language model Mistral-7b. It is divided into two subsets corresponding to context-rich and context-scarce scenarios, with 100,000 and 300,000 samples respectively. This dataset is utilized for fine-tuning pre-trained encoder models to improve their performance across various tasks.
EmoPillars 数据集概述
基本信息
- 许可证: Apache-2.0
- 任务类别: 文本分类
- 语言: 英语 (en)
- 数据集名称: EmoPillars
- 数据规模: 100K < n < 1M
数据集描述
- 内容: 包含28个类别的细粒度无上下文和上下文情感分类的合成数据。
- 生成方法: 使用多步流程基于Mistral模型生成。
- 用途: 用于训练多标签分类器,识别28种情感类别的话语,可选择在给定情境(上下文)中识别。
相关资源
- 分类器集合: https://huggingface.co/collections/alex-shvets/emopillars-67ec9694541e0bc69d62861f
- GitHub仓库: https://github.com/alex-shvets/emopillars
- 论文: https://arxiv.org/abs/2504.16856
引用信息
bibtex @misc{shvets2025emopillarsknowledgedistillation, title={Emo Pillars: Knowledge Distillation to Support Fine-Grained Context-Aware and Context-Less Emotion Classification}, author={Alexander Shvets}, year={2025}, eprint={2504.16856}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2504.16856} }




