SEARCHGEN-20K, SEARCHGEN-BENCH, SEARCHGEN-CORPUS-1M
收藏资源简介:
SEARCHGEN-20K是由香港科技大学、滑铁卢大学等机构联合构建的大规模、双语多模态数据集,旨在系统研究视觉生成中的世界知识瓶颈问题。该数据集包含20,839条涵盖12种失败类别和22个领域的文本到图像提示,平均每条提示包含5.2个知识缺口,并标注了34,694个视觉参考槽和16,345个文本知识槽,数据来源于真实用户请求和种子实体库的合成生成。其构建过程采用基于模板的人工构思与LLM辅助重写策略,并预先执行了145,642次搜索会话以形成可复现的SEARCHGEN-CORPUS-1M语料库。该数据集主要应用于代理增强视觉生成、搜索策略优化和知识边界发现等领域,致力于解决生成模型在开放世界知识场景下的幻觉与评估盲区问题。
SEARCHGEN-20K is a large-scale bilingual multimodal dataset jointly constructed by The Hong Kong University of Science and Technology, the University of Waterloo and other institutions, aiming to systematically investigate the world knowledge bottleneck problem in visual generation. This dataset contains 20,839 text-to-image prompts covering 12 failure categories and 22 domains, with an average of 5.2 knowledge gaps per prompt. It is annotated with 34,694 visual reference slots and 16,345 text knowledge slots, and its data originates from real user requests and synthetic content generated from seed entity libraries. Its construction adopts a template-based manual conceptualization and LLM-assisted rewriting strategy, and pre-conducted 145,642 search sessions to form a reproducible SEARCHGEN-CORPUS-1M corpus. This dataset is primarily applied in fields such as agent-augmented visual generation, search strategy optimization, and knowledge boundary discovery, and aims to address the hallucination and evaluation blind spot problems of generative models in open-world knowledge scenarios.





