遇见数据集

Kenpache/multilingual-financial-sentiment

收藏
Hugging Face2026-04-10 更新2026-04-12 收录
官方服务:

资源简介:

--- license: apache-2.0 language: - en - zh - ja - de - fr - es - ar tags: - finance - sentiment-analysis - multilingual - financial-news - text-classification size_categories: - 10K<n<100K task_categories: - text-classification task_ids: - sentiment-classification --- # Multilingual Financial Sentiment Dataset A curated dataset of **39,829** financial news sentences annotated with sentiment labels (Negative / Neutral / Positive) across **7 languages**, collected from **80+ financial news sources** worldwide. ## Dataset Summary | | | |---|---| | **Total samples** | 39,829 | | **Languages** | 7 (EN, ZH, JA, DE, FR, ES, AR) | | **Labels** | 3 (negative, neutral, positive) | | **Format** | CSV | | **Sources** | 80+ financial news outlets | ## Languages | Language | Code | Samples | % of Total | |---|---|---|---| | Japanese | ja | 8,287 | 20.8% | | Chinese | zh | 7,930 | 19.9% | | Spanish | es | 7,125 | 17.9% | | English | en | 6,887 | 17.3% | | German | de | 5,023 | 12.6% | | French | fr | 3,935 | 9.9% | | Arabic | ar | 642 | 1.6% | ## Label Distribution ### Overall | Label | Count | % | |---|---|---| | Neutral | 18,130 | 45.5% | | Positive | 12,257 | 30.8% | | Negative | 9,442 | 23.7% | ### Per Language | Language | Negative | Neutral | Positive | |---|---|---|---| | Japanese | 1,767 | 3,376 | 3,144 | | Chinese | 1,921 | 3,126 | 2,883 | | Spanish | 1,641 | 3,842 | 1,642 | | English | 1,704 | 3,339 | 1,844 | | German | 1,392 | 2,425 | 1,206 | | French | 870 | 1,657 | 1,408 | | Arabic | 147 | 365 | 130 | ## Sources Data was collected from major financial news outlets across all target languages: - **English:** CNBC, Yahoo Finance, Fortune, Bloomberg, Reuters, Barron's, Benzinga, Seeking Alpha, Kiplinger, Business Insider, Moneycontrol, Zacks, FT - **Chinese:** Sina Finance, EastMoney, 10jqka, NBD, China Securities, 163 Finance, Hexun, STCN - **Japanese:** Nikkan Kogyo, Nikkei, Reuters JP, Minkabu, Asahi Business, ZUU Online, Toyo Keizai, ITmedia Business, Sankei Economy - **German:** Börse.de, NTV Börse, FAZ Finanzen, Wallstreet Online, Börse Online, OnVista, Manager Magazin, Tagesschau, Handelsblatt, WiWo, Süddeutsche - **French:** Boursorama, Tradingsat, BFM Business, Le Revenu, L'Expansion, Capital, Le Figaro Bourse, L'AGEFI, EasyBourse - **Spanish:** Estrategias de Inversión, Expansión, El Confidencial, Cinco Días, Bloomberg Línea, Investing.es, Bolsamanía, EFE Economía, Infobae Economía, El Financiero, Portafolio, DF.cl - **Arabic:** Al Khaleej Economy, Al Jazeera Economy, Sabq Economy, RT Arabic Economy, Okaz Economy, Sky News Business, Maaal, Al Arabiya Economy, CNBC Arabia ## Data Format CSV with 4 columns: | Column | Type | Description | |---|---|---| | `sentence` | string | Financial news text | | `label` | string | Sentiment: `negative`, `neutral`, or `positive` | | `source` | string | News source identifier | | `language` | string | ISO 639-1 language code | ## Loading the Dataset ```python from datasets import load_dataset dataset = load_dataset("Kenpache/multilingual-financial-sentiment") df = dataset["train"].to_pandas() # Filter by language en_data = df[df["language"] == "en"] # Filter by label positive = df[df["label"] == "positive"] ``` ### Direct CSV Loading ```python import pandas as pd df = pd.read_csv("hf://datasets/Kenpache/multilingual-financial-sentiment/all_languages_clean.csv") print(df.head()) ``` ## Sample Data | sentence | label | source | language | |---|---|---|---| | Revenue surged 40% year-over-year, beating expectations. | positive | yahoo_finance | en | | Die Aktie verlor nach der Gewinnwarnung deutlich an Wert. | negative | faz_finanzen | de | | 同社の業績は前年並みで推移している。 | neutral | nikkei | ja | | Les résultats du groupe sont conformes aux attentes. | neutral | boursorama | fr | ## License & Usage This dataset is released **for academic and non-commercial research only** under fair use / text and data mining exceptions (EU DSM Directive Art. 3). Each sample is a single short sentence (≤1-2 sentences) extracted for sentiment classification research. Copyright of original texts remains with their respective publishers, cited in the `source` field. **For commercial use**, you must obtain licenses from original sources directly. If you are a rights holder and want content removed, open an issue or email alisterclrouli@gmail.com — we will remove within 48h. ## License Apache 2.0

--- 许可证:Apache 2.0 语言: - 英语(en) - 中文(zh) - 日语(ja) - 德语(de) - 法语(fr) - 西班牙语(es) - 阿拉伯语(ar) 标签: - 金融(finance) - 情感分析(sentiment-analysis) - 多语言(multilingual) - 金融新闻(financial-news) - 文本分类(text-classification) 规模类别: - 10K<n<100K(10000至100000条样本) 任务类别: - 文本分类(text-classification) 任务子项: - 情感分类(sentiment-classification) --- # 多语言金融情感数据集 本数据集为经过精心整理的39829条金融新闻语句数据集,标注了情感标签(负面/中性/正面),覆盖7种语言,数据源自全球80余家金融新闻媒体。 ## 数据集概览 | | | |---|---| | **总样本数** | 39,829 | | **覆盖语言** | 7种(EN、ZH、JA、DE、FR、ES、AR) | | **标签类别** | 3种(负面、中性、正面) | | **数据格式** | CSV | | **数据来源** | 80余家金融新闻媒体 | ## 覆盖语言 | 语言 | 代码 | 样本数 | 占总样本比例 | |---|---|---|---| | 日语 | ja | 8,287 | 20.8% | | 中文 | zh | 7,930 | 19.9% | | 西班牙语 | es | 7,125 | 17.9% | | 英语 | en | 6,887 | 17.3% | | 德语 | de | 5,023 | 12.6% | | 法语 | fr | 3,935 | 9.9% | | 阿拉伯语 | ar | 642 | 1.6% | ## 标签分布 ### 整体分布 | 标签 | 数量 | 占比 | |---|---|---| | 中性 | 18,130 | 45.5% | | 正面 | 12,257 | 30.8% | | 负面 | 9,442 | 23.7% | ### 分语言分布 | 语言 | 负面 | 中性 | 正面 | |---|---|---|---| | 日语 | 1,767 | 3,376 | 3,144 | | 中文 | 1,921 | 3,126 | 2,883 | | 西班牙语 | 1,641 | 3,842 | 1,642 | | 英语 | 1,704 | 3,339 | 1,844 | | 德语 | 1,392 | 2,425 | 1,206 | | 法语 | 870 | 1,657 | 1,408 | | 阿拉伯语 | 147 | 365 | 130 | ## 数据来源 数据采集自覆盖所有目标语言的主流金融新闻媒体: - **英语**:CNBC、Yahoo Finance(雅虎财经)、Fortune(《财富》)、Bloomberg(彭博社)、Reuters(路透社)、Barron's(巴伦周刊)、Benzinga、Seeking Alpha、Kiplinger、Business Insider、Moneycontrol、Zacks、FT(英国《金融时报》) - **中文**:新浪财经(Sina Finance)、东方财富(EastMoney)、同花顺(10jqka)、每日经济新闻(NBD)、中国证券报(China Securities)、网易财经(163 Finance)、和讯网(Hexun)、证券时报网(STCN) - **日语**:Nikkan Kogyo、Nikkei(日经新闻)、Reuters JP(路透社日本版)、Minkabu、Asahi Business(朝日商业)、ZUU Online、Toyo Keizai(东洋经济)、ITmedia Business、Sankei Economy(产经经济) - **德语**:Börse.de、NTV Börse、FAZ Finanzen、Wallstreet Online、Börse Online、OnVista、Manager Magazin(《经理人杂志》)、Tagesschau(今日新闻)、Handelsblatt(《商报》)、WiWo、Süddeutsche(《南德意志报》) - **法语**:Boursorama、Tradingsat、BFM Business、Le Revenu、L'Expansion(《扩张》)、Capital(《资本》)、Le Figaro Bourse(费加罗财经)、L'AGEFI、EasyBourse - **西班牙语**:Estrategias de Inversión、Expansión(《扩张报》)、El Confidencial(《机密报》)、Cinco Días(《五天报》)、Bloomberg Línea、Investing.es、Bolsamanía、EFE Economía(埃菲社财经)、Infobae Economía、El Financiero(《金融家报》)、Portafolio、DF.cl - **阿拉伯语**:Al Khaleej Economy、Al Jazeera Economy(半岛电视台财经)、Sabq Economy、RT Arabic Economy(今日俄罗斯阿拉伯语财经)、Okaz Economy、Sky News Business(天空新闻财经)、Maaal、Al Arabiya Economy(阿拉伯电视台财经)、CNBC Arabia(CNBC阿拉伯频道) ## 数据格式 CSV格式,包含4列: | 列名 | 数据类型 | 描述 | |---|---|---| | `sentence` | 字符串 | 金融新闻文本 | | `label` | 字符串 | 情感标签:`negative`(负面)、`neutral`(中性)或`positive`(正面) | | `source` | 字符串 | 新闻来源标识 | | `language` | 字符串 | ISO 639-1语言代码 | ## 数据集加载 python from datasets import load_dataset dataset = load_dataset("Kenpache/multilingual-financial-sentiment") df = dataset["train"].to_pandas() # 按语言过滤数据 en_data = df[df["language"] == "en"] # 按标签过滤数据 positive = df[df["label"] == "positive"] ### 直接加载CSV文件 python import pandas as pd df = pd.read_csv("hf://datasets/Kenpache/multilingual-financial-sentiment/all_languages_clean.csv") print(df.head()) ## 样本示例 | 金融新闻语句 | 情感标签 | 新闻来源 | 语言代码 | |---|---|---|---| | 营收同比飙升40%,超出市场预期。 | positive | yahoo_finance | en | | 该公司发布盈利预警后,股价大幅下跌。 | negative | faz_finanzen | de | | 该公司业绩与上年持平。 | neutral | nikkei | ja | | 该集团业绩符合市场预期。 | neutral | boursorama | fr | ## 许可证与使用规范 本数据集仅可用于学术研究与非商业用途,符合合理使用/文本与数据挖掘豁免条款(欧盟数字单一市场指令第3条)。 每条样本均为单条短语句(1-2句以内),专为情感分类研究提取。原始文本的版权归各自发布方所有,相关信息已在`source`字段中注明。 **如需商业使用**,请直接与原始新闻来源获取授权许可。 若您为内容版权方并希望移除相关数据,请提交Issue或发送邮件至alisterclrouli@gmail.com,我们将在48小时内处理。 ## 许可证 Apache 2.0

提供机构:
Kenpache
二维码
社区交流群
二维码
科研交流群
商业服务