遇见数据集

YADHU1234/dataset_1M

收藏
Hugging Face2025-02-24 更新2025-04-12 收录
官方服务:

资源简介:

该数据集包含两个文本特征:印地语(hindi)和马拉雅拉姆语(malayalam)。数据集分为训练集,共有100万个示例,总大小约为264MB。提供了默认配置,用于指定训练数据文件的路径。

The dataset includes two text features: Hindi (hindi) and Malayalam (malayalam). It is split into a training set with a total of 1,000,000 examples, totaling approximately 264MB in size. A default configuration is provided to specify the path to the training data files.

提供机构:
YADHU1234
二维码
社区交流群
二维码
科研交流群
商业服务