HPMD
收藏资源简介:
HPMD(Historical Persian Manuscript Dataset)是帕亚姆·努尔大学创建的历史波斯手稿数据集,专为单词定位任务设计。该数据集包含223页、3678行、37631个单词和130630个字符,源自多本历史波斯诗歌与散文书籍,并由11名标注者在区域、行和文本级别进行注释,未标注单词级边界以降低标注成本。创建过程精心选取不同流派、字体和布局的页面,通过Transkribus平台完成三层标注,并附有页面级属性标签。数据集主要服务于历史手稿的单词定位研究,旨在解决缺乏公开历史波斯手稿数据集及单词级标注昂贵的问题,支持仅需行级标注即可实现单词搜索的方法,助力历史学家快速定位特定名称、日期或事件。
HPMD (Historical Persian Manuscript Dataset) is a historical Persian manuscript dataset developed by Payame Noor University, specifically designed for word localization tasks. This dataset consists of 223 pages, 3678 lines, 37,631 words and 130,630 characters, sourced from multiple historical Persian poetry and prose books. It was annotated by 11 annotators at the region, line and text levels, while word-level boundaries were not labeled to reduce annotation costs. During its development, pages of different genres, scripts and layouts were carefully selected. Three-level annotation was completed via the Transkribus platform, with page-level attribute tags attached. This dataset primarily serves research on word localization for historical manuscripts, aiming to address the issues of the lack of publicly available historical Persian manuscript datasets and the high cost of word-level annotation. It supports methods that enable word spotting with only line-level annotations, assisting historians in quickly locating specific names, dates or events.

- 1HPMD: A Historical Persian Manuscript Dataset for Word Spotting with Line-Level Annotation帕亚姆·努尔大学 · 2026年



