SynthGlyph Dataset, DesignText Dataset
收藏资源简介:
SynthGlyph Dataset是由北京大学王选计算机研究所构建的大规模合成字符数据集,包含4194种TrueType字体渲染的6857个字符,总计约2880万条实例,支持中英文字符及符号的多样化风格迁移。DesignText Dataset则聚焦真实设计场景,收录11.55万条设计样本,涵盖背景图、文本描述及细粒度标注,通过自动化流程整合多源数据。两数据集采用合成渲染与真实标注相结合的方法构建,旨在解决图形设计中风格化文本编辑与生成的难题,为AI辅助平面设计提供高精度训练资源。
The SynthGlyph Dataset is a large-scale synthetic character dataset constructed by the Wangxuan Institute of Computer Technology, Peking University. It contains 6,857 characters rendered with 4,194 TrueType fonts, totaling approximately 28.8 million instances, and supports diverse style transfer for Chinese and English characters as well as symbols. The DesignText Dataset focuses on real-world design scenarios, collecting 115,500 design samples covering background images, text descriptions and fine-grained annotations, and integrates multi-source data via automated workflows. Both datasets are constructed by combining synthetic rendering and real annotation methods, aiming to address the challenges of stylized text editing and generation in graphic design, and provide high-precision training resources for AI-assisted graphic design.



