TextSculpt-Data
收藏资源简介:
TextSculpt-Data是由北京交通大学和字节跳动联合构建的大规模高保真场景文本编辑数据集,旨在解决现有开源训练数据稀缺和评估标准缺失的问题。该数据集包含120万条经过OCR验证的文本到图像样本以及200万对源-目标图像对齐、背景一致性强的文本编辑配对样本,总计320万条训练数据,数据来源结合了基于VLM的标题改写、高质量图像合成以及程序化文本渲染与合成技术。其构建过程采用自动化流水线,通过程序化渲染确保文本编辑的精确性和背景区域的严格保留。该数据集主要应用于场景文本编辑领域,支持文本添加、替换、删除和混合编辑等核心任务,旨在提升模型在文本渲染和编辑方面的精确性与视觉真实性。
TextSculpt-Data is a large-scale high-fidelity scene text editing dataset jointly constructed by Beijing Jiaotong University and ByteDance, aiming to address the issues of scarce open-source training data and lack of standardized evaluation metrics. This dataset contains 1.2 million OCR-validated text-to-image samples and 2 million pairs of source-target aligned text editing paired samples with strong background consistency, totaling 3.2 million training samples. Its data sources integrate VLM-based caption rewriting, high-quality image synthesis, and procedural text rendering and synthesis technologies. The dataset is built via an automated pipeline, which ensures the accuracy of text editing and strict retention of background regions through procedural rendering. This dataset is primarily applied in the field of scene text editing, supporting core tasks such as text addition, replacement, deletion and hybrid editing, aiming to enhance the accuracy and visual realism of models in text rendering and editing.
数据集名称:TextSculptor: Training and Benchmarking Scene Text Editing
数据集地址:
- 数据集下载页面:https://huggingface.co/datasets/dafbgd/TextSculpt-Data
数据集状态:
- 已发布:TextSculpt-Data(训练/微调用数据集)
- 待发布:TextSculpt-Bench(评测基准)及评测脚本
相关资源:
- 论文预印本:https://arxiv.org/abs/2605.21090
引用格式: bibtex @article{lin2026textsculptor, title={TextSculptor: Training and Benchmarking Scene Text Editing}, author={Lin, Yiheng and Jiao, Siyu and Lan, Xiaohan and Zhou, Wei and She, Qi and Yu, Fei and Chen, Heyun and Wang, Zhengwei and Chen, Jinghuan and Li, Moran and Yu, Yingchen and Feng, Zijian and Zhao, Yao and Wei, Yunchao and Zhong, Yujie}, journal={arXiv preprint arXiv:2605.21090}, year={2026} }
说明:该数据集用于场景文本编辑(Scene Text Editing)任务的训练与基准测试,目前仅开放了训练数据部分(TextSculpt-Data),评测集(TextSculpt-Bench)和评估代码尚未发布。




