SAIL-Caption
收藏资源简介:
SAIL-Caption是由字节跳动抖音内容组创建的大规模视觉理解数据集,包含1亿张图像样本,旨在为视觉语言模型(VLM)提供高质量的预训练数据。该数据集通过多任务、多节点、多处理的异步标注系统生成,确保了数据的多样性和高质量。数据集的内容涵盖了丰富的视觉元素,如独特的n-gram、名词、动词和形容词,显著优于其他公开的标注数据集。SAIL-Caption的创建过程包括数据收集、参考数据标注、标注模型训练和大规模数据生成。该数据集主要用于视觉语言模型的预训练,旨在提升模型在视觉理解和指令跟随任务中的表现。
SAIL-Caption is a large-scale visual understanding dataset created by the Douyin Content Team of ByteDance, which contains 100 million image samples and aims to provide high-quality pre-training data for Vision-Language Models (VLMs). This dataset is generated through an asynchronous annotation system with multi-task, multi-node and multi-processing capabilities, ensuring the diversity and high quality of the data. The dataset covers a wide range of visual elements, such as distinctive n-grams, nouns, verbs and adjectives, and significantly outperforms other publicly available annotated datasets. The creation process of SAIL-Caption includes data collection, reference data annotation, annotation model training and large-scale data generation. This dataset is primarily used for the pre-training of vision-language models, with the objective of enhancing the model's performance in visual understanding and instruction-following tasks.

- 1Scalable Vision Language Model Training via High Quality Data Curation字节跳动抖音内容组 · 2025年



