caveman273/aida-handwritten
收藏资源简介:
该数据集包含来自AIDA项目的手写文本行图像及其转录文本。它是完整AIDA数据集的一个子集,仅包含**质量最佳的手写**注释——即注释者对每个字符都确信无误的文本行。大多数文本行是芬兰语,也有一些瑞典语、英语、法语和德语的文本行。数据集主要用于手写文本识别(HTR)任务。数据集的每一行包含:`image`(文本行图像,PNG格式)、`text`(转录文本)和`file_name`(原始图像文件名)。数据集分为训练集(6943条)、验证集(1151条)和测试集(1270条)。数据来源于芬兰商业中央档案馆(ELKA),包括信件、船舶记录、商业出版物等各种文档类型。原始文本的生产者包括个人和不同公司的员工。注释工作由芬兰国家档案馆和ELKA的员工完成。此外,还通过合成数据增加了训练数据量,使用了TextRecognitionDataGenerator库和来自古登堡计划及互联网档案馆的芬兰书籍和杂志。数据集未匿名化,可能包含个人姓名等敏感信息。
This dataset contains handwritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the **best-quality handwritten** annotations — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German. The dataset was created for handwritten text recognition (HTR). Each row contains: `image` (the textline image in PNG format), `text` (the ground-truth transcription), and `file_name` (the original image filename). The dataset is split into train (6943 lines), validation (1151 lines), and test (1270 lines). The data is collected from Central Archives for Finnish Business (ELKA) and consists of various document types including letters, ship records, and business publications. The original texts were produced by private individuals and employees of different companies. Annotations were done by employees of National Archives of Finland and ELKA. Synthetic data was also generated using the TextRecognitionDataGenerator library and Finnish books from Project Gutenberg and magazines from the Internet Archive. The dataset is not anonymized and may contain personal names and other sensitive information.



