遇见数据集

ardauzunoglu/dclm_random200m

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

该数据集包含197,212个训练样本,总大小约为12.1GB,下载大小约为734.7MB。每个样本包含以下字段:sample_index(样本索引,int64类型)、scanned_index(扫描索引,int64类型)、text(文本内容,字符串类型)、source_dataset(来源数据集名称,字符串类型)和url(来源链接,字符串类型)。数据仅提供训练分割,文件路径为data/train-*。

This dataset contains 197,212 training samples with a total size of approximately 12.1GB and a download size of approximately 734.7MB. Each sample includes the following fields: sample_index (int64), scanned_index (int64), text (string), source_dataset (string), and url (string). The data is provided only in a training split, with file paths as data/train-*.

提供机构:
ardauzunoglu
二维码
社区交流群
二维码
科研交流群
商业服务