OUHD-L: Online Urdu Handwriting Dataset – Lines
收藏资源简介:
OUHD-L (Online Urdu Handwriting Dataset - Lines) is the first publicly available online Urdu handwritten text-line dataset. It contains 2,403 handwritten text lines from 311 writers, recorded using a Wacom IoT paper device at 500 Hz. Each sample captures per-stroke metadata including pen pressure, tilt, velocity, and stroke order. Writers. Participants ranged in age from 17 to 35 years (mean 19.7) and comprised 86 female (27.7%) and 225 male (72.3%) writers, predominantly university-age students. (At the sample level this is 686 female / 1,717 male of the 2,403 lines.) Splits. Writers are partitioned with no overlap across splits: Train (1,911 samples, 248 writers), Val (239 samples, 31 writers), and Test (253 samples, 32 writers). Split manifests are provided as train.csv, val.csv, and test.csv. Files. Each sample includes a raw pen-trajectory CSV (28 columns: x/y coordinates, pressure, tilt, height, thickness, timestamp, and 20 derived kinematic features), a rendered handwriting image, and 23 grayscale feature visualisation images (one per kinematic channel: speed, acceleration, curvature, stroke timing, etc.). Total file counts: 2,403 trajectory CSVs, 2,403 main images, 55,269 feature images. The dataset was built to fill a gap in line-level online corpora for cursive Nastaliq Urdu. It is well-suited for online handwriting recognition, multimodal fusion of online and offline representations, and sequence modelling for right-to-left scripts. Data collection was carried out at NUST and NCAI, Islamabad, Pakistan. This dataset accompanies the paper "Online Urdu Text-Line Recognition by Bridging Stroke Dynamics and Offline Representations", accepted at ICDAR 2026.



