A Curated Bangladesh-Based Dataset of Handwritten and Printed Prescription Images
收藏资源简介:
This dataset comprises 200 de-identified prescription images collected in Bangladesh, including printed and handwritten prescriptions written in Bangla, English, and mixed Bangla-English formats. Direct identifiers such as patient and physician names, phone numbers, physician degrees, signatures, chamber details, and registration numbers have been removed for privacy protection. The dataset is distributed as a ZIP archive containing an images folder with the prescription images, a labels folder with corresponding text files containing YOLO-style bounding box coordinates for medicine-name regions, and a CSV file listing the image number and the number of annotated medicine boxes for each prescription image. The dataset is intended for research and educational use in optical character recognition, handwriting recognition, prescription parsing, multilingual medical document understanding, object detection, and healthcare document analysis.
本数据集囊括200张采集自孟加拉国的去标识化处方图像,涵盖孟加拉语、英语及孟英混合格式的打印版与手写版处方。为保护隐私,已移除患者与医师姓名、电话号码、医师资质、签名、诊所信息及注册编号等直接标识符。本数据集以ZIP压缩包形式分发,内含存储处方图像的images文件夹、存储对应文本文件的labels文件夹(文本文件包含药品名称区域的YOLO风格边界框坐标),以及一份列明每张处方图像的图像编号与标注药品框数量的CSV文件。 本数据集可应用于光学字符识别、手写识别、处方解析、多语言医疗文档理解、目标检测及医疗文档分析领域的研究与教学工作。




