TWIG-50K
收藏资源简介:
TWIG-50K是由香港中文大学与美团联合构建的大规模多模态数据集,专为视觉生成中的文本推理任务设计。该数据集包含5万条高质量样本,涵盖丰富的文本-视觉对齐数据,源自人工标注与自动化流程的结合。其核心应用于增强生成模型在复杂场景下的语义控制能力,通过监督微调有效解决视觉幻觉与指令遵循问题,推动交互式视觉合成技术的发展。
TWIG-50K is a large-scale multimodal dataset jointly developed by The Chinese University of Hong Kong and Meituan, specifically designed for text reasoning tasks in visual generation. This dataset contains 50,000 high-quality samples covering rich text-vision aligned data, which is derived from a combination of manual annotation and automated workflows. Its core application is to enhance the semantic control ability of generative models in complex scenarios, effectively addressing visual hallucination and instruction following issues via supervised fine-tuning, and promoting the development of interactive visual synthesis technologies.
Thinking-while-Generating (TwiG) 数据集概述
数据集基本信息
- 数据集名称: Thinking-while-Generating (TwiG)
- 数据集规模: TwiG-50K
- 官方存储库: https://github.com/ZiyuGuo99/Thinking-while-Generating
- 关联论文: "Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation"
核心创新
- 首创框架: 首个将文本推理交织在整个视觉合成过程中的框架
- 方法特点: 在生成过程中直接将文本思考编织到展开的画布中
- 技术优势: 在生成过程中提供实时语义指导和反思
框架架构
TwiG将生成过程解耦为三个核心模块:
- 调度模块: 决定何时思考
- 推理模块: 决定说什么
- 反思模块: 决定如何优化
可视化展示
- 定性比较: 展示与其他方法的对比效果
- 反思能力: 展示框架的反思优化能力
- 思考过程: 展示完整的思考生成过程
引用信息
bibtex @article{guo2026thinking, title={Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation}, author={Guo, Ziyu and Zhang, Renrui and Li, Hongyu and Zhang, Manyuan and Chen, Xinyan and Wang, Sifan and Feng, Yan and Pei, Peng and Heng, Pheng-Ann}, journal={arXiv:2511.16671}, year={2025} }

- 1Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation香港中文大学, 美团 · 2025年



