遇见数据集

BCD-Dataset

收藏
Zenodo2026-08-26 更新2026-10-01 收录
官方服务:

资源简介:

The BCD-Dataset (Brazilian Clinical Documents Dataset) is a challenging dataset developed for research on Handwritten Text Recognition (HTR) in real-world Brazilian clinical documents. The dataset contains 13,916 segmented text-line images written in Brazilian Portuguese and extracted from clinical forms used in healthcare-related workflows. Unlike conventional handwriting benchmarks, BCD-Dataset includes substantial variability in handwriting styles, textual content, field structure, abbreviations, numerical information, and domain-specific terminology, reflecting the complexity of documents encountered in practical clinical environments. For reproducible experimentation, the dataset is provided with predefined stratified splits containing 11,132 training samples, 1,391 validation samples, and 1,393 test samples. BCD-Dataset was designed to support the development and evaluation of handwriting recognition systems under challenging real-world conditions, enabling research on model robustness, domain-specific HTR, semantic-category performance, and the recognition of heterogeneous handwritten clinical information. The dataset accompanies the paper “BCD-Dataset: a Challenge Dataset of Clinical Documents for Handwritten Text Recognition,” accepted for publication at SIBGRAPI 2026 – Conference on Graphics, Patterns and Images.

提供机构:
Zenodo
创建时间:
2026-08-26
二维码
社区交流群
二维码
科研交流群
商业服务