遇见数据集

nmixx-fin/opensource_korean_finance_datasets

收藏
Hugging Face2025-01-15 更新2025-02-15 收录
官方服务:

资源简介:

这是一个包含金融相关文本数据的数据集,其中包含文本来源(source)、文本内容(text)、文本类别(category)和词汇计数(token_count)四个字段。数据集以韩语为主要语言,并提供训练集(train)分割。训练集大小为532344662字节,共有502831个示例。

This dataset contains finance-related text data, including four fields: text source (source), text content (text), text category (category), and word count (token_count). The dataset is primarily in Korean and provides a training set (train) split. The training set is 532344662 bytes in size and contains 502831 examples.

提供机构:
nmixx-fin
二维码
社区交流群
二维码
科研交流群
商业服务