ICFHR 2020 Competition on Image Retrieval for Historical Handwritten Fragments (HisFrag20) Dataset
收藏资源简介:
This competition investigates the performance of large-scale retrieval of historical document fragments based on writer recognition. The analysis of historic fragments is a difficult challenge commonly solved by trained humanists.<br> We focus on the task of automatic image retrieval to simulate common scenarios of humanities research, such as fragment or writer retrieval. Therefore, we created a large dataset consisting of more than 120000 fragments.<br> The goal is then to find similar patches of the same page or manuscript. contains ~100 000 fragments using the Historical-IR19 as base dataset, they should all contain some text, however some fragments are quite small. Training-set: contains ~100 000 fragments using the Historical-IR19 as base dataset, they should all contain some text, however some fragments are quite small. Test-set: contains about 20 000 new fragments Naming-convention: WID_PID_FID.jpg , where WID=writer id, PID: page id, FID= fragment id For more information visit: https://lme.tf.fau.de/research/competitions/hisfragir20/
本竞赛旨在探究基于书写者识别的历史文献残片大规模检索性能。历史残片的分析是一项极具挑战性的任务,目前通常由经过专业训练的人文领域学者完成。<br>我们聚焦于自动图像检索任务,以模拟人文研究中的常见应用场景,例如残片检索或书写者检索。为此,我们构建了一个包含逾12万条残片的大型数据集。<br>本次任务的目标为检索得到同一页面或手稿的相似残片。本数据集以Historical-IR19为基础数据集,包含约10万条残片,所有残片均包含文本,但部分残片尺寸极小。训练集(Training-set):以Historical-IR19为基础数据集,包含约10万条残片,所有残片均包含文本,但部分残片尺寸极小。测试集(Test-set):包含约2万条全新残片。命名规范为:WID_PID_FID.jpg,其中WID为书写者ID,PID为页面ID,FID为残片ID。更多相关信息请访问:https://lme.tf.fau.de/research/competitions/hisfragir20/



