Hunyuan-Recap100M
收藏资源简介:
Hunyuan-Recap100M是一个由腾讯Hunyuan Team发布的合成字幕数据集,包含1亿个图像-文本对。该数据集通过直接偏好优化方法(DPO)生成,显著降低了hallucinations(虚假信息)的出现,同时富含知识信息,旨在解决大规模视觉语言模型预训练中的数据不足问题。数据集的构建目的是为视觉语言模型预训练、跨模态生成等研究领域提供高质量的多模态数据。
Hunyuan-Recap100M is a synthetic caption dataset released by the Tencent Hunyuan Team, which contains 100 million image-text pairs. Generated via the Direct Preference Optimization (DPO) method, this dataset significantly reduces the incidence of hallucinations while being abundant in knowledge. It is designed to address the issue of insufficient training data for large-scale vision-language model pre-training. The dataset is constructed to provide high-quality multimodal data for research fields including vision-language model pre-training and cross-modal generation.

- 1Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training腾讯 · 2025年



