BID
收藏资源简介:
该存储库引入了名为巴西身份证件数据集(BID 数据集)的数据集:巴西身份证件的第一个公共数据集。 BID 数据集在工作中提出:“BID 数据集:文档处理任务的挑战数据集”,旨在解决计算机视觉领域的三个关键挑战:(i)文档图像分类; (ii) 文本区域分割和 (iii) 光学字符识别 (OCR)。 BID Dataset 由巴西身份证件的图像组成,分为八类:国家驾驶执照(CNH)的正面和背面、CNH 正面、CNH 背面、自然人登记(CPF)正面、CPF 背面、通用登记(RG) 正面、RG 背面和 RG 正面和背面。 BID 数据集由 28,800 张文档图像组成,每个类别有 3,600 个样本。
This repository introduces the Brazilian Identity Document Dataset (BID Dataset), the first public dataset focused on Brazilian identity documents. Proposed in the work entitled "BID Dataset: A Challenging Dataset for Document Processing Tasks", the BID Dataset aims to address three core challenges in the field of computer vision: (i) document image classification; (ii) text region segmentation; and (iii) optical character recognition (OCR). The BID Dataset consists of images of Brazilian identity documents, categorized into eight classes: front and back sides of the National Driver's License (CNH), front of CNH, back of CNH, front of the Natural Person Registration (CPF), back of CPF, front of General Registration (RG), back of RG, and both front and back sides of RG. The BID Dataset includes a total of 28,800 document images, with 3,600 samples per category.

- BID数据集首次发表于《生物信息学》杂志,标志着该数据集的正式诞生。
- BID数据集首次应用于蛋白质相互作用网络的分析,展示了其在生物信息学领域的潜力。
- BID数据集被广泛应用于多个生物医学研究项目,成为研究蛋白质功能和相互作用的重要工具。
- BID数据集的更新版本发布,增加了新的蛋白质相互作用数据,提升了数据集的完整性和准确性。
- BID数据集被纳入多个国际生物信息学数据库,进一步扩大了其影响力和应用范围。



