Emuru训练数据集
收藏资源简介:
Emuru训练数据集是一个大规模的合成数据集,由摩德纳和雷焦艾米利亚大学的研究团队创建。该数据集包含220万张文本图像,这些图像由不同的背景和超过10万种字体渲染的英文文本线组成。数据集中的文本内容均匀分布,涵盖了多种英语语料库中的词汇,旨在训练Emuru模型,使其能够生成不含有背景噪声、风格多样的手写文本图像,以用于文档分析和图形设计等领域。
The Emuru Training Dataset is a large-scale synthetic dataset created by a research team from the University of Modena and Reggio Emilia. This dataset contains 2.2 million text images, which are composed of English text lines rendered with various backgrounds and over 100,000 fonts. The text content in the dataset is uniformly distributed, covering vocabulary from multiple English corpora. It is designed to train the Emuru model to generate handwritten text images with diverse styles and free of background noise, which can be applied to fields such as document analysis and graphic design.

- 1Zero-Shot Styled Text Image Generation, but Make It Autoregressive摩德纳和雷焦艾米利亚大学 · 2025年



