commoncatalog-cc-by-recap-qwen3p5-35b-a3b
收藏资源简介:
该数据集是 CommonCatalog CC-BY 数据集的重新描述版本,由 Qwen3.5-35B-A3B 模型生成。它包含 14,576,560 条生成的描述文本,对应 14,576,558 个图像资产,不包含图像数据。数据集中的每一行通过 Flickr 的 photoid(表示为 asset_instance_id)以及源 Parquet 文件的行定位器(image_shard 和 image_member)与上游的 common-canvas/commoncatalog-cc-by 数据集对齐。除了本数据集新生成的 caption_text 字段外,还保留了上游数据集中的原始字段:caption、description、title、usertags、machinetags 以及 blip2_caption。这些上游字段不属于本数据集的输出,本数据集也未对其进行修改。描述生成采用 Qwen3.5-35B-A3B-FP8 模型,单阶段长描述,输出限制 512 token,使用 Qwen3.5 发布时的默认采样参数,并通过 vLLM 0.17.1 提供服务。禁用思考模式,输出仅进行空白字符清理,未进行语义后处理。该数据集仅供研究和审计可重复性使用,不适合直接用于训练:用户在使用前必须进行语义后处理、安全与策略过滤、URL/内容身份验证以及图像权利审查。CC BY 4.0 许可证仅适用于本数据集中生成的描述文本,不适用于源图像或第三方元数据。
This dataset is a re-captioning version of the CommonCatalog CC-BY dataset, generated by the Qwen3.5-35B-A3B model. It contains 14,576,560 generated caption texts corresponding to 14,576,558 image assets, without including image data. Each row is aligned to the upstream common-canvas/commoncatalog-cc-by dataset via Flickrs photoid (denoted as asset_instance_id) and row locators (image_shard and image_member) from the source Parquet file. In addition to the newly generated caption_text field, it retains upstream original fields: caption, description, title, usertags, machinetags, and blip2_caption. These upstream fields are not part of this datasets output and have not been modified. The description generation uses the Qwen3.5-35B-A3B-FP8 model with single-stage long descriptions, output limit of 512 tokens, default sampling parameters from Qwen3.5 release, served via vLLM 0.17.1. Thinking mode is disabled, output is only whitespace-cleaned without semantic post-processing. This dataset is intended for research and reproducibility auditing only, not suitable for direct training: users must perform semantic post-processing, safety and policy filtering, URL/content authentication, and image rights review before use. The CC BY 4.0 license applies only to the generated description texts in this dataset, not to source images or third-party metadata.
数据集概述
数据集名称:CommonCatalog CC-BY recaptions with Qwen3.5-35B-A3B
数据集地址:https://huggingface.co/datasets/BootsofLagrangian/commoncatalog-cc-by-recap-qwen3p5-35b-a3b
数据集内容
- 规模:包含 14,576,560 条生成的标题(captions),对应 14,576,558 个图像资源。
- 数据格式:仅包含标题文本(caption-only),不包含图像文件。每个条目包括:
asset_instance_id(来源于Flickrphotoid)image_shard和image_member(源Parquet文件定位器)caption_text(由Qwen3.5模型生成的标题)- URL和内容哈希值,用于独立核对。
- 来源:与公开的
common-canvas/commoncatalog-cc-by数据集在修订版本80f50fe4a1ca937f37a11be3f8eee5199d776ff3下匹配。
生成方法
- 模型:
Qwen/Qwen3.5-35B-A3B-FP8 - 生成设置:
- 单阶段长标题生成
- 输出限制为512个token
- 使用Qwen3.5发布时的采样默认值
- 自由格式输出
- 推理服务:vLLM 0.17.1
- 禁用思考模式(thinking disabled)
- 仅去除空格,无语义后处理
- 配置文件:见
generation_config.yaml
使用说明
- 加载方式:可通过
load_dataset流式加载标题数据,或从源数据集加载配对的图像数据(示例代码见README)。 - 注意:源数据集的文本字段(如
caption、description、title、usertags、machinetags、blip2_caption)属于上游数据集,非本仓库生成内容。 - 发布目的:用于研究和审计复现。不适用于直接训练,用户需自行进行语义后处理、安全与政策过滤、URL/内容身份验证以及图像权利审查。
许可与引用
-
许可协议:
CC BY 4.0,仅适用于本仓库生成的标题文本,不涵盖源图像或第三方元数据。 -
引用文献:
@misc{oh2026matchedbudget, title={A Matched-Budget Audit Framework for Recaptioned Image-Text Supervision Distributions}, author={Giyeong Oh and Junghun Park and Yuhan Bae and Youngjae Yu}, year={2026} }




