HGAF: Handwritten German Administrative Forms
收藏资源简介:
HGAF: Handwritten German Administrative Forms Version: 1.0.0Authors: Tobias Kalmbach, Luis Bast, Andreas WagnerContact: tobias.kalmbach@proton.meLicense: CC BY 4.0 Summary HGAF contains 291 scanned handwritten forms: 97 forms in each of three layouts (A, B, C). All layouts have the same set of fields, but present them in different structures: Layout A: a simple form with large, clearly delimited fields and ample writing space. Designed as the easiest case. Layout B: a compact, tabular form with small fields, designed to constrain the writing space. Layout C: a form without clearly delimited fields, leaving placement and size of the handwriting to the participant. Each participant filled in one form per layout with identical content. The three forms of one participant therefore differ only in visual layout. This makes it possible to test cross-layout generalization while keeping the writer and the content constant. Two participants filled out two sets of forms, each with a different ground truth and a different script (one in cursive, one in block letters). All field contents are synthetic. The ground truth was generated with Faker and an LLM, and participants signed with fake signatures. The field contents therefore contain no real personal data. The handwriting images themselves were released with the participants' consent (see Data collection). No identifiers for the participants were saved. Intended use: field-level handwritten text recognition (HTR) and form understanding, including cross-layout generalization.We suggest splitting by doc_id and keeping all three layouts of a doc_id in the same split. Note that two participants contributed two form sets each under different doc_ids, so doc_id-disjoint splits are not fully writer-disjoint. Contents Every scan is provided at three resolutions: 100, 300 and 600 dpi. Each resolution was scanned separately. HGAF/ ├── 100dpi/ │ ├── layoutA/ │ │ ├── 100dpi-A-001.png … 100dpi-A-097.png │ │ └── GroundTruth_A.json │ ├── layoutB/ (same structure, GroundTruth_B.json) │ └── layoutC/ (same structure, GroundTruth_C.json) ├── 300dpi/ (same structure as 100dpi/) ├── 600dpi/ (same structure as 100dpi/) ├── empty_examples/ blank forms (A, B, C) and an example instruction sheet (doc_id 1) └── README.md The number in the file name is its document ID (doc_id) with leading zeros.The ground truth is identical across resolutions. Each resolution folder contains a copy of it so that every folder can be used on its own. Ground-truth format Each GroundTruth_<layout>.json contains one entry per form: Key Description file image file name, including extension doc_id form identifier. It is also printed in the upper-right corner of the form, next to the layout letter. The same doc_id identifies the same participant and content across all three layouts. index ordinal position of the file within its folder ground_truth object with one value per field (12 fields, see below) First entry of GroundTruth_A.json for 300 dpi, truncated: { "file": "300dpi-A-001.png", "doc_id": "1", "index": 0, "ground_truth": { "last_name": "Textor", "first_name": "Linda", "monthly_salary": "256,99", "…": "…" } } All values are stored as UTF-8 strings. Umlauts and ß are kept as written. Numeric-looking fields (postal code, phone number, house number, income) should be read as strings, because leading zeros and formatting are part of the ground truth. Fields Label on form (German) Ground truth key Notes Name last_name Vorname first_name Monatliches Einkommen monthly_salary Geburtsdatum birthday E-Mail-Adresse email Telefonnummer phone Studiengang studies one of four degree programmes: IM, WIN, BWL, Kunst Straße street Hausnummer house_number PLZ postal_code Ort city Kommentar comment Unterschrift — signature not transcribed, not in the ground truth Generation of field contents About 70% of the values were generated with Faker (German locale de_DE) and about 30% with Microsoft Copilot (an LLM). The LLM was used to add non-German names, addresses and formats. All values were then revised manually to make sure the data contains enough variation and edge cases. Specific characteristics per field: Names: include non-German names. Monthly income: integers and decimals, mostly in German notation (comma as decimal separator). Birthday: mostly DD.MM.YYYY. Some dates use - or / as separators, and some spell out the month (DD. Monat YYYY). Implausible dates are included on purpose: dates in the future (relative to the generation date, 08.05.2026) and very old dates (e.g. the year 1200). E-mail: only reserved example domains (example.com, example.org, …). The addresses do not necessarily match the name fields. Phone numbers: German and international formats. Address (street, house number, postal code, city): the fields are generated independently. The combination does not necessarily correspond to a real address. Studies: evenly distributed across the four degree programmes. Every form has exactly one box ticked. Comment: entirely LLM-generated. Mostly German, with some English, Spanish and French. Signature: participants were asked to use a fake signature. Data collection Participants: 95 writers, mostly university students. No personal or demographic data was collected. Procedure: each participant received three forms (layouts A, B, C) together with a pre-generated record and copied the record by hand into all three forms. Digitization: all forms were scanned at 100, 300 and 600 dpi on the same scanner. Consent: participants gave informed consent to the public release of the scanned forms, including their handwriting. Ground-truth verification: after collection, every field of every form was compared with the pre-generated record. Wherever the handwriting differed from the record (typos, omissions, abbreviations), the ground truth was corrected to match what is actually written on the form. The ground truth therefore reflects the ink, not the original record. Despite this check, residual errors cannot be ruled out. Known limitations Small and homogeneous writer population (mostly students), single collection site. Content was copied rather than written freely, so the handwriting may be more careful than in real administrative use. Synthetic content: value distributions (e.g. names, addresses, implausible dates) do not reflect real-world frequencies. Residual ground-truth errors are possible (see Data collection). License The dataset is released under CC BY 4.0. How to cite @dataset{kalmbach_hgaf_2026, author = {Kalmbach, Tobias and Bast, Luis and Wagner, Andreas}, title = {{HGAF}: Handwritten German Administrative Forms}, year = {2026}, publisher = {Zenodo}, version = {1.0.0}, doi = {10.5281/zenodo.22934701} } Changelog 1.0.0 2026-09-24: initial release.



