ExpertVerse-100K
收藏资源简介:
ExpertVerse-100K是由阿里巴巴集团与清华大学联合构建的、面向知识密集型视觉合成任务的大规模基准数据集。该数据集包含约10万条高质量样本,每条数据均配备逐步推理链(CoT)和知识锚定的原理注释,其数据来源融合了人工精心设计的种子提示与专有视觉语言模型(VLM)的大规模合成扩展。数据集通过分层扩展策略构建,覆盖了8个专家领域和9种认知能力划分出的58个子类别,确保了知识密度与任务多样性。该数据集的核心应用在于推动下一代生成式模型在跨学科、多跳知识推理方面的能力评估与训练,旨在解决现有模型在专业领域知识深度理解与复杂视觉逻辑生成方面的瓶颈问题。
ExpertVerse-100K is a large-scale benchmark dataset jointly constructed by Alibaba Group and Tsinghua University, targeting knowledge-intensive visual synthesis tasks. This dataset contains approximately 100,000 high-quality samples, each equipped with step-by-step Chain-of-Thought (CoT) reasoning chains and knowledge-anchored principle annotations. Its data sources combine manually curated seed prompts and large-scale synthetic expansion from proprietary Vision-Language Models (VLM). The dataset is constructed via a hierarchical expansion strategy, covering 58 subcategories derived from 8 expert domains and 9 cognitive capabilities, ensuring knowledge density and task diversity. The core application of this dataset lies in promoting the capability evaluation and training of next-generation generative models in cross-disciplinary, multi-hop knowledge reasoning, aiming to address the bottleneck issues faced by existing models in deep understanding of professional domain knowledge and complex visual logic generation.
数据集概述:ExpertVerse
- 数据集名称:ExpertVerse
- 核心定位:一个面向专家级推理的通用基准测试,专注于知识密集型视觉合成场景。
- 构建目的:评估当前多模态生成模型在知识密集型生成任务上的推理能力,因为现有方法多局限于显式常识推理、浅层因果理解和直接知识召回。
- 数据规模:包含 1,611 个专家标注的实例。
- 任务类型:覆盖三种视觉生成任务:
- 单图像编辑
- 多图像合成
- 文本到图像生成
- 分类体系:基于正交的分类学结构,涵盖:
- 9 种认知能力
- 8 个专家学科
- 细分为 58 个子学科
- 衍生数据集:基于 ExpertVerse,通过自动化流程构建了大规模的 ExpertVerse-100K 数据集,包含推理轨迹和知识锚定的理由注释。
- 关联模型:基于 ExpertVerse 数据集训练了名为 KnowThinker 的视觉语言模型推理引擎,该引擎采用强化学习微调,能够联合生成思考过程和精细化指令,并使用了定制的 Bootstrapped Pareto Policy Optimization (BPPO) 优化方法。

- 1ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis阿里巴巴集团; 清华大学 · 2026年




