TextPecker
收藏资源简介:
TextPecker是由华中科技大学与字节跳动联合构建的视觉文本渲染(VTR)结构异常标注数据集,旨在解决生成图像中文本结构失真、模糊、错位等细粒度缺陷的感知问题。该数据集包含真实生成模型产出的文本图像(含字符级结构异常标注)以及通过笔画编辑引擎合成的扩充数据,覆盖中英双语场景,数据来源包括TextAtlas5M、Lex-10k等文本语料及多模态生成模型(如Stable Diffusion、Qwen-Image)。其构建过程通过人工精细标注与合成引擎增强相结合,显著提升了结构错误类型的多样性。该数据集为文本生成模型的强化学习优化提供了细粒度结构感知能力,推动高保真视觉文本渲染技术的发展。
TextPecker is a visual text rendering (VTR) structural anomaly annotation dataset jointly developed by Huazhong University of Science and Technology and ByteDance. It aims to address the perceptual problems of fine-grained defects such as text structural distortion, blurriness and misalignment in generated images. This dataset contains text images produced by real generative models (with character-level structural anomaly annotations) and augmented data synthesized via a stroke editing engine, covering both Chinese and English scenarios. Its data sources include text corpora such as TextAtlas5M and Lex-10k, as well as multimodal generative models like Stable Diffusion and Qwen-Image. The construction process combines manual fine-grained annotation and synthetic engine augmentation, which significantly enhances the diversity of structural error types. This dataset provides fine-grained structural perception capabilities for the reinforcement learning optimization of text generative models, and promotes the advancement of high-fidelity visual text rendering technology.
- 1TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering华中科技大学; 字节跳动 · 2026年



