KhmerST
收藏资源简介:
KhmerST数据集是由拉罗谢尔大学的信息图像交互实验室创建的,专门用于低资源高棉语场景文本检测和识别任务。该数据集包含1,544张专家标注的图像,涵盖室内和室外场景,具有多样化的文本类型和光照条件。数据集的创建过程包括在柬埔寨各地采集图像,并使用VGG图像标注器进行标注,提供多边形边界框坐标和行级文本信息。KhmerST数据集旨在解决高棉语在自然场景中的文本检测和识别问题,适用于数字文档存档、自动翻译服务和增强技术应用的可达性。
The KhmerST dataset was developed by the Laboratory of Information, Image and Interaction of the University of La Rochelle, specifically tailored for low-resource Khmer natural scene text detection and recognition tasks. It comprises 1,544 expert-annotated images covering both indoor and outdoor scenarios, featuring diverse text types and lighting conditions. The dataset construction workflow includes collecting images across Cambodia and annotating them using the VGG Image Annotator, providing polygon bounding box coordinates and line-level text information. The KhmerST dataset aims to address the challenges of Khmer text detection and recognition in natural scenes, and supports applications such as digital document archiving, automatic translation services, and accessibility enhancement of technical applications.




