遇见数据集

Punjab Pakistan Synthetic and Real License Plate Character Datasets (P-LPCD)

收藏
Zenodo2025-09-23 更新2026-05-26 收录
官方服务:

资源简介:

Dataset Overview There is limited availability of license plate datasets specific to Punjab, Pakistan. With evolving plate templates and privacy concerns around capturing real vehicle images, collecting large-scale real data is challenging. A practical approach is to gather a modest set of real images and fine-tune models pretrained on synthetic license plates, enabling adaptation to new templates without much effort. The Punjab Pakistan Synthetic and Real License Plate Character Datasets (P-LPCD) are designed for license plate character detection and recognition, focusing exclusively on vehicle plates from Punjab, Pakistan. The dataset includes two complementary subsets: PS-LPCD (Punjab Synthetic License Plate Character Dataset), containing 40,000 images (32,000 for training and 8,000 for validation) and intended for large-scale pretraining on synthetic images, and PR-LPCD (Punjab Real License Plate Character Dataset), containing 650 images (500 for training and 150 for testing) and intended for fine-tuning and evaluation on real-world images. Subset Number of Images Split Purpose PS-LPCD 40,000 Train: 32,000Validation: 8,000 Large-scale pretraining on synthetic images PR-LPCD 650 Train: 500Test: 150 Fine-tuning and evaluation on real-world images Key Features Covers four official Punjab license plate templates (2 front, 2 rear) to ensure realistic design variations. Labels are provided in YOLO format (.txt) with bounding boxes for individual characters. Contains 36 classes: Alphabets: A–Z Digits: 0–9 The special “PUNJAB” class: identifies regions on the license plate that are not part of the standard character sequence. These regions appear as “PUNJAB” on the plate, can be ignored during evaluation or post-processing, and any characters misidentified within them can also be skipped. Recommended Workflow Pretrain models on the PS-LPCD (synthetic) subset. Fine-tune models on the PR-LPCD (real) subset. Evaluate using the PR-LPCD test set for real-world performance. Applications License plate character detection and recognition Synthetic-to-real domain adaptation Benchmarking OCR pipelines for regional license plates Notes This dataset is not global; it is specifically tailored to the fonts, templates, and styles of Punjab, Pakistan license plates. The code used to generate the synthetic data, along with additional supporting files, is available in the GitHub repository P-LPCD and is included when the dataset is downloaded.

提供机构:
Zenodo
创建时间:
2025-09-23
二维码
社区交流群
二维码
科研交流群
商业服务