Word-level controlled scene text dataset
收藏资源简介:
WordCon数据集是一个用于场景文本渲染的词汇级控制数据集,由香港理工大学、新加坡国立大学、重庆大学、浙江大学和Tiamat AI的研究人员构建。该数据集旨在解决场景文本渲染中单词级别的字体样式控制问题。数据集包含场景文本渲染所需的词汇级控制数据,以及每个单词的分割掩码。数据集的创建过程涉及文本-图像对齐框架(TIA),它利用了 grounding 模型的跨模态对应关系来增强文本到图像模型的训练。WordCon数据集的应用领域包括艺术文本渲染、文本编辑和条件图像文本渲染,旨在解决文本渲染中单词级别的样式控制问题。
The WordCon dataset is a vocabulary-level controlled dataset for scene text rendering, constructed by researchers from The Hong Kong Polytechnic University, National University of Singapore, Chongqing University, Zhejiang University, and Tiamat AI. This dataset aims to address the problem of word-level font style control in scene text rendering. The dataset includes vocabulary-level control data required for scene text rendering, as well as segmentation masks for each individual word. The creation of this dataset involves the Text-Image Alignment (TIA) framework, which leverages the cross-modal correspondence of grounding models to enhance the training of text-to-image models. The application scenarios of the WordCon dataset cover artistic text rendering, text editing, and conditional image text rendering, with the goal of solving the problem of word-level style control in text rendering.

- 1WordCon: Word-level Typography Control in Scene Text RenderingThe Hong Kong Polytechnic University, National University of Singapore, Chongqing Univesity, Zhejiang University, Tiamat AI · 2025年



