遇见数据集

AbstractPhil/cc-prompts-sharded

收藏
Hugging Face2026-05-15 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是Conceptual Captions数据集的一个分片版本,专门为Qwen任务转换流程准备。原始数据经过去重处理,仅保留至少4个词的标题,并分割为三个大致相等的分片,以及一个long分片,用于包含超过50个词的长标题(保留供后续高容量模型处理)。每个数据行包括id(在CC流中的位置,用8位数字零填充)、caption(文本标题)和n_words(词数)。id在多次运行中保持稳定,用于下游推理中的恢复跟踪。数据集适用于文本生成任务,语言为英语,大小在100万到1000万样本之间,许可证为CC-BY-4.0。

This dataset is a sharded version of the Conceptual Captions dataset, prepared for the Qwen task-conversion pipeline. The original data has been deduplicated, filtered to captions with a minimum of 4 words, and split into three roughly-equal shards plus a long shard for captions over 50 words (reserved for later high-capacity model processing). Each row contains an id (the position in the CC stream, zero-padded to 8 digits), a caption (text), and n_words (word count). The id is stable across re-runs and used for resume tracking in downstream inference. The dataset is intended for text-generation tasks, in English, with a size between 1M and 10M samples, licensed under CC-BY-4.0.

提供机构:
AbstractPhil
二维码
社区交流群
二维码
科研交流群
商业服务