ro_sft_pixmo_cap
收藏资源简介:
PixmoCap-RO是PixmoCap数据集的罗马尼亚语翻译版本。PixmoCap原始数据集包含非常长(平均约200词)且详细的图像描述。本数据集使用Seed-X-PPO模型进行翻译,旨在为罗马尼亚语视觉语言模型提供高质量的指令微调数据。它是论文《Înțelegi românește? A Recipe for Romanian Vision-Language Models》中提出的罗马尼亚语VLM指令微调方案的关键组成部分。数据集包含约60.8万个训练样本,总大小约193.6 GB。每个样本包含图像序列、指令、描述文本以及多轮对话格式的消息记录。数据集适用于罗马尼亚语视觉语言理解、图像描述生成、多模态对话等任务的模型训练与评估。
PixmoCap-RO is the Romanian translation version of the PixmoCap dataset. The original PixmoCap dataset contains very long (average about 200 words) and detailed image descriptions. This dataset is translated using the Seed-X-PPO model and aims to provide high-quality instruction fine-tuning data for Romanian vision-language models. It is a key component of the Romanian VLM instruction fine-tuning scheme proposed in the paper Înțelegi românește? A Recipe for Romanian Vision-Language Models. The dataset contains approximately 608,000 training samples, with a total size of about 193.6 GB. Each sample includes image sequences, instructions, description texts, and message records in a multi-turn dialogue format. The dataset is suitable for model training and evaluation in tasks such as Romanian vision-language understanding, image description generation, and multimodal dialogue.
数据集概述:ro_sft_pixmo_cap
该数据集是罗马尼亚语版本的 PixmoCap 数据集,专为罗马尼亚视觉语言模型(VLM)的指令微调而设计。
- 许可证: cc-by-nc-4.0
- 语言: 罗马尼亚语(ro)
- 起源: 基于 PixmoCap 数据集(包含平均约200词的长篇详细描述),通过 Seed-X-PPO 模型翻译为罗马尼亚语。
- 用途: 作为 "Înțelegi românește?" A Recipe for Romanian Vision-Language Models(Masala 等人,2026年,arXiv:2605.31401)中提出的罗马尼亚VLM指令微调协议的一部分。
- 数据集大小: 总大小约 193.6 GB,下载大小约 189.3 GB。
- 划分: 仅包含训练集(train),共 608,026 个样本。
数据特征
每条数据包含以下字段:
- id (字符串): 样本的唯一标识符。
- images (图像序列): 与样本关联的图像。
- instruction (字符串): 指令文本。
- caption (字符串): 图像的描述文本。
- messages (列表): 对话消息序列,每条消息包含:
- role (字符串): 消息角色。
- content (字符串): 消息内容。
引用
bibtex @inproceedings{deitke2025molmo, title={Molmo and pixmo: Open weights and open data for state-of-the-art vision-language models}, author={Deitke, Matt and Clark, Christopher and Lee, Sangho and Tripathi, Rohun and Yang, Yue and Park, Jae Sung and Salehi, Mohammadreza and Muennighoff, Niklas and Lo, Kyle and Soldaini, Luca and others}, booktitle={Proceedings of the Computer Vision and Pattern Recognition Conference}, pages={91--104}, year={2025} }
bibtext @misc{masala2026intelegi, title={``^{I}nc{t}elegi Rom^{a}nec{s}te? A Recipe for Romanian Vision-Language Models}, author={Mihai Masala and Marius Leordeanu and Mihai Dascalu and Traian Rebedea}, year={2026}, eprint={2605.31401}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2605.31401}, }




