multilingual-financial-sentiment
收藏资源简介:
多语言金融情感数据集是一个精心整理的金融新闻句子数据集,包含39,829条标注了情感标签(负面/中性/正面)的句子,涵盖7种语言(英语、中文、日语、德语、法语、西班牙语、阿拉伯语)。数据来自全球80多家金融新闻媒体。数据集以CSV格式提供,包含四个字段:句子文本、情感标签、新闻来源和语言代码。各语言样本分布为:日语20.8%、中文19.9%、西班牙语17.9%、英语17.3%、德语12.6%、法语9.9%、阿拉伯语1.6%。整体情感标签分布为:中性45.5%、正面30.8%、负面23.7%。该数据集适用于多语言情感分析研究,特别是金融领域的文本分类任务。数据集采用Apache 2.0许可,仅限学术和非商业研究使用。
The Multilingual Financial Sentiment Dataset is a meticulously curated collection of financial news sentences, containing 39,829 labeled entries with sentiment tags (negative, neutral, positive) across 7 languages: English, Chinese, Japanese, German, French, Spanish, and Arabic. The data is sourced from over 80 global financial news media outlets. The dataset is distributed in CSV format, with four fields: sentence text, sentiment label, news source, and language code. The sample distribution per language is: Japanese 20.8%, Chinese 19.9%, Spanish 17.9%, English 17.3%, German 12.6%, French 9.9%, and Arabic 1.6%. The overall sentiment label distribution is: neutral 45.5%, positive 30.8%, and negative 23.7%. This dataset is applicable to multilingual sentiment analysis research, particularly text classification tasks in the financial domain. The dataset is licensed under Apache 2.0, and is restricted to academic and non-commercial research use only.




