遇见数据集

fpadovani/goldfish-Dp-english-10mb-tokenized

收藏
Hugging Face2026-05-19 更新2026-05-31 收录
官方服务:

资源简介:

这是一个用于自然语言处理(NLP)任务的数据集,包含训练分割(train),共有2,977,689个示例。数据集特征包括input_ids(int32列表,表示文本的编码标识符)和attention_mask(int8列表,用于注意力机制中的掩码)。数据集总大小约为1.08GB,下载大小约为1.72GB,适用于模型训练和文本处理应用。

This is a dataset for natural language processing (NLP) tasks, containing a training split (train) with 2,977,689 examples. The dataset features include input_ids (list of int32, representing encoded identifiers for text) and attention_mask (list of int8, used for masking in attention mechanisms). The total dataset size is approximately 1.08GB, with a download size of about 1.72GB, suitable for model training and text processing applications.

提供机构:
fpadovani
二维码
社区交流群
二维码
科研交流群
商业服务