遇见数据集

A Multi-Session Smartphone-Scanned Handwriting Corpus for Open-Set Writer Identification

收藏
Mendeley Data2026-08-04 收录
官方服务:

资源简介:

Smartphone-scanned handwritten laboratory records from a first-year engineering Makerspace course: 65 writers x 17 lab experiments, 565 PDF records, 2,463 pages at 300 DPI. Every page was captured by the students themselves with a mobile-phone camera through consumer scanning apps (Adobe Scan, CamScanner; app watermarks appear on some pages) — no flatbed scanner anywhere in the corpus. The data is multi-session (each experiment written and scanned on a different day), diagram-rich (inline circuit and mechanical sketches, pencil shading), double-sided cursive Indian-college English — a capture condition and demographic absent from existing writer-identification benchmarks such as IAM, CVL, CERUG and Firemaker. CONTENTS. corpus/ms_<experiment>.zip — 17 zips, one per lab experiment, each containing student<NNN>.pdf records. tables/dataset_long.csv — one row per PDF (student_id, experiment, filename, page count). tables/dataset_matrix.csv — student x experiment page-count matrix. tables/dataset_summary.json — corpus totals and processing flags. tables/splits.json — the canonical writer-disjoint train/validation/test split (45/10/10) for reproducible benchmarking. checksums_sha256.txt — SHA-256 of every file. NAMING AND ANONYMISATION. Every record is student<NNN>.pdf, where NNN is a 3-digit pseudonym assigned per writer; original upload filenames are never included and no mapping to real identities is published. Valid ids: student002–student136, even numbers only. Experiment identity is defined by the folder name. Exactly one PDF per student per experiment; byte-identical re-uploads were removed (MD5-verified) and duplicate scans resolved by page count. KNOWN ARTIFACTS (documented and deliberately retained): mirrored show-through and unmirrored stack-transparency ghosts from thin double-sided paper, ruled lines, uncontrolled lighting and perspective. ETHICS AND USAGE. Handwriting is a biometric. Released for non-commercial research in document analysis and writer identification (CC BY-NC 4.0). Do not attempt to re-identify writers; do not use this data to train handwriting-forgery systems targeting real individuals.

某工科一年级创客空间(Makerspace)课程的智能手机扫描手写实验记录:涵盖65名书写者、17项实验,共565份PDF记录,总计2463页,分辨率为300 DPI。所有页面均由学生本人使用手机摄像头配合大众消费级扫描应用(Adobe Scan、CamScanner;部分页面带有应用水印)采集完成,本数据集未使用任何平板扫描仪。 该数据集为多会话场景(每项实验均在不同日期书写并扫描),包含大量图表(内嵌电路与机械草图、铅笔阴影),采用印度高校英文手写草书且为双面书写——此类采集条件与人群分布,在现有手写者识别(writer identification)基准数据集(如IAM、CVL、CERUG及Firemaker)中均未出现。 【数据集内容】 corpus/ms_<experiment>.zip:共17个压缩包,对应17项实验,每个压缩包内包含student<NNN>.pdf格式的记录文件。 tables/dataset_long.csv:每一行对应一份PDF记录,字段包括student_id(学生ID)、experiment(实验编号)、filename(文件名)、page count(页数)。 tables/dataset_matrix.csv:学生×实验的页数矩阵。 tables/dataset_summary.json:数据集总统计信息与处理标记。 tables/splits.json:用于可复现基准测试的标准手写者分离式训练/验证/测试划分(比例为45/10/10)。 checksums_sha256.txt:所有文件的SHA-256校验值。 【命名与匿名化处理】 所有记录均采用student<NNN>.pdf命名格式,其中NNN为分配给每位书写者的3位数字化名;原始上传文件名未被保留,且未公开任何与真实身份的映射关系。有效ID范围为student002–student136,仅包含偶数编号。实验身份由文件夹名称定义。每位学生对应每项实验仅一份PDF记录;经MD5校验后,移除了字节完全一致的重复上传文件,并通过页数解决了重复扫描的问题。 【已知采集伪影(已记录并刻意保留)】 薄型双面纸张导致的镜像透印与非镜像堆叠透印虚影、页面格线、光照与透视角度不受控的图像问题。 【伦理与使用规范】 手写笔迹属于生物特征信息。本数据集仅面向非商业研究发布,可用于文档分析与手写者识别研究(授权协议为CC BY-NC 4.0)。请勿尝试重新识别书写者身份;请勿使用本数据集训练针对真实个体的手写伪造系统。

创建时间:
2026-07-17
二维码
社区交流群
二维码
科研交流群
商业服务