遇见数据集

fpadovani/goldfish-Dp-indonesian-10mb-tokenized

收藏
Hugging Face2026-05-23 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个基于文本处理的机器学习数据集,包含训练集分割,共4,147,487个示例,总大小为1,185,532,766字节。特征包括input_ids(int32列表)和attention_mask(int8列表),可能用于自然语言处理任务,如序列分类或文本生成。数据文件存储在data/train-*路径下,下载大小为1,878,707,525字节。具体用途和内容来源未在README中说明。

This dataset is a machine learning dataset for text processing, containing a training split with 4,147,487 examples and a total size of 1,185,532,766 bytes. Features include input_ids (list of int32) and attention_mask (list of int8), likely used for natural language processing tasks such as sequence classification or text generation. Data files are stored in the path data/train-*, with a download size of 1,878,707,525 bytes. Specific use cases and content sources are not detailed in the README.

提供机构:
fpadovani
二维码
社区交流群
二维码
科研交流群
商业服务