LehongWu/example-rewritten-collect_3tasks_25A_v3_1_0423-gemini3flash_medium-repeat8_suc200trajs
收藏资源简介:
该数据集是一个多模态交互数据集,包含图像和文本提示对,用于训练或评估视觉-语言模型。数据集中每个样本包括图像列表、多轮对话提示(含内容和角色信息)、奖励模型信息(含真实标签和风格)、额外信息(如答案、完成状态、思考过程、唯一标识符、目标、任务特定提示和先前指令),以及数据来源、能力描述和划分信息。数据集仅提供测试集,共789个样本,适用于多模态任务的研究和评估。
This dataset is a multimodal interaction dataset containing image and text prompt pairs, designed for training or evaluating vision-language models. Each sample includes a list of images, multi-turn prompts (with content and role information), reward model information (including ground truth and style), extra information (such as answer, completion status, thought process, UUID, goal, task-specific prompt, and previous instruction), as well as data source, ability description, and split information. The dataset only provides a test set with 789 samples, suitable for research and evaluation of multimodal tasks.




