RuoliuYang/ULVR_v2_premium
收藏资源简介:
ULVR_v2_premium是一个用于通用潜在视觉推理(ULVR)的高质量精选子集,源自ULVR_v2_clean数据集。它包含262,556个样本,每个样本都经过严格的质量评估和筛选,确保视觉推理步骤的有效性和必要性。数据集按推理类别分为8个独立分割,每个分割针对不同的视觉推理任务:bbox_highlight(高亮感兴趣区域后回答)、bbox_crop(裁剪感兴趣区域后回答,与bbox_highlight共享相同样本但形式不同)、text_cot(基于图像的文本链式推理)、helper_interleaved(在推理轨迹中交错插入辅助图像)、visual_representation(生成/使用抽象表示如深度/边缘/分割)、scene_graph(关系场景图推理)、chart_focus(聚焦/缩放图表区域后回答)和doc_crop(裁剪文档区域后回答)。样本选择基于质量层级,优先包含辅助救援(P1)和两者都正确(P2)的样本,排除低质量项。数据集仅用于训练,不包含验证分割,适用于多模态视觉问答和视觉推理任务。
ULVR_v2_premium is a high-quality curated subset for Universal Latent Visual Reasoning (ULVR), derived from the ULVR_v2_clean dataset. It contains 262,556 samples, each of which has undergone rigorous quality assessment and filtering to ensure the validity and necessity of visual reasoning steps. The dataset is divided into 8 independent splits based on reasoning categories, each targeting a distinct visual reasoning task: bbox_highlight (answer queries after highlighting regions of interest), bbox_crop (answer queries after cropping regions of interest, sharing identical samples with bbox_highlight but in different formats), text_cot (text chain-of-thought reasoning based on images), helper_interleaved (interleaving auxiliary images within the reasoning trajectory), visual_representation (generating or utilizing abstract representations such as depth, edge, and segmentation), scene_graph (relational scene graph reasoning), chart_focus (answer queries after focusing or zooming in on chart regions), and doc_crop (answer queries after cropping document regions). Sample selection follows a quality tiering principle, prioritizing samples from Primary Rescue (P1) and Both Correct (P2) while excluding low-quality items. The dataset is intended solely for training purposes, without a validation split included, and is suitable for multimodal visual question answering and visual reasoning tasks.




