遇见数据集

The ICDAR 2003 Informal Competition for the Recognition of On-line Words: The Unipen-ICROW-03 benchmark set - Version 0.0

收藏
Zenodo2023-02-10 更新2026-05-26 收录
数据链接:
官方服务:

资源简介:

Proposal for an informal benchmark on word recognition. See for the related ImUnipen collection<br> of word images from on-line vectorial handwriting data: https://zenodo.org/record/1195059 At the time (ICDAR 2003) there was not a lot of interest so the project was not pursued. Lambert Schomaker - February 2023 _______________________________________________________________________________ The ICDAR 2003 Informal Competition for the Recognition of On-line Words:<br> The Unipen-ICROW-03 benchmark set <br> Version 0.0 Lambert Schomaker / International Unipen Foundation The ICROW suite of test files for the recognition of isolated on-line<br> free-style (handprint, mixed and cursive) words has been<br> composed. Different tablets, nationalities and languages<br> are involved. Only the ASCII set is used within word labels. The set contains: 13119 written words<br> 884 unique lexical word entries<br> 72 writers Language: Dutch, English, Italian.<br> Nationalities: Dutch, Irish, Italian, + mixed The benchmark test is a good estimator for <br> "walk-up" recognition performance. [Note: some of the writers (NIC-Pc95*.dat set) are present in the<br> UNIPEN R01/V07 distribution, but the actual words are unseen <br> outside of the Int. Unipen Foundation.] Please note the Copyright notice in the <br> accompanying file 'Copyright' Wed Jul 16 21:20:10 CEST 2003 Lambert Schomaker --------------------------------------------------------------------------- Instructions for the ICDAR 2003 informal competition for<br> the recognition of on-line words. 1 - unpack the .tgz file<br> 2 - use the UNIPEN files as input for your recognizer.<br> 3 - report, for each writer, a file &lt;writer-id&gt;.res Example: do-my-recognizer &lt; NIC-Hi93b-marc.dat &gt; NIC-Hi93b-marc.res Format of the .res file. No XML for this moment: simplicity does it. We assume that the recognizer is able to produce a top-10 list<br> of likely words, sorted from most likely to least likely.<br> The output for each word is on a single line. The correct<br> target word is in the first column. &lt;targetword 1&gt; &lt;best word hyp.&gt; &lt;2nd-best word hyp.&gt; ... &lt;10th-best word hyp&gt;<br> &lt;targetword 2&gt; &lt;best word hyp.&gt; &lt;2nd-best word hyp.&gt; ... &lt;10th-best word hyp&gt; Example with two words: summertime slumbertime slipknot summertime somatome spumante simulative semitone schoolmate sermonette semimature<br> Aberdeen Adamson Aberdeen Addison Armageddon Abyssinian Araban Albanian Alabamian Abraham Adelaide <br> 4 - pack the *.res files in a .tgz or .zip file and send them<br> to schomaker@ai.rug.nl<br> All *.dat files need to be processed. LS.<br>

提供机构:
Zenodo
创建时间:
2023-02-10
二维码
社区交流群
二维码
科研交流群
商业服务