ButterChicken98/soyabean_unified_rag_short_ablation
收藏资源简介:
该数据集是一个多模态数据集,包含图像和文本字段,主要用于图像描述、标签分类和提示生成任务。数据集包含2800个训练示例,每个示例包括图像、标题、多个标题变体(captions)、标签以及多种提示字段,如随机提示、基于检索的提示(k1)和基于检索增强生成(RAG)的提示(k2_slm)。这些提示字段涉及ID、分数、方面(aspect)和回退(fallback)信息,支持自然语言处理和计算机视觉的研究与应用。数据集还包含RAG相关的短类键(rag_short_class_key)和源语料库(rag_short_source_corpus)字段。整体结构旨在促进图像-文本交互、内容生成和检索任务的实验。
This dataset is a multimodal dataset containing image and text fields, primarily designed for image captioning, label classification, and prompt generation tasks. It consists of 2800 training examples, each including an image, a caption, multiple caption variants (captions), a label, and various prompt fields such as random prompts, retrieval-based prompts (k1), and retrieval-augmented generation (RAG) prompts (k2_slm). These prompt fields involve IDs, scores, aspects, and fallback information, supporting research and applications in natural language processing and computer vision. The dataset also includes RAG-related fields like rag_short_class_key and rag_short_source_corpus. The overall structure aims to facilitate experiments in image-text interaction, content generation, and retrieval tasks.



