amalia-llm/PorTEXTO
收藏资源简介:
PorTEXTO是第一个针对当代和文化相关的欧洲葡萄牙语(pt-PT)视觉文本提取的基准数据集。它专注于现代、真实世界的葡萄牙语内容,包括手写笔记、野外场景文本和合成图像,为OCR(光学字符识别)和大型视觉语言模型提供了一个具有挑战性的评估套件。数据集包含四个配置:handwritten(从扫描文档中裁剪的手写文本区域,351个样本)、handwritten_full_page(全页扫描的手写文档,121个样本)、synthetic(为pt-PT OCR评估合成的文本图像,200个样本)和in_the_wild(在自然环境中捕获的真实世界场景文本,107个样本),总样本数为779。每个样本包括图像、标准pt-PT转录提示和真实转录。数据集采用cc-by-nc-4.0许可,仅限非商业用途,并需注明出处。
PorTEXTO is the first benchmark dataset for visual text extraction targeting contemporary and culturally relevant European Portuguese (pt-PT). It focuses on modern, real-world Portuguese content, including handwritten notes, in-the-wild scene text, and synthetic images, acting as a challenging evaluation suite for Optical Character Recognition (OCR) and Large Vision-Language Models (LVLMs). The dataset includes four configurations: handwritten (handwritten text regions cropped from scanned documents, 351 samples), handwritten_full_page (full-page scanned handwritten documents, 121 samples), synthetic (text images synthesized for pt-PT OCR evaluation, 200 samples), and in_the_wild (real-world scene text captured in natural environments, 107 samples), with a total of 779 samples. Each sample consists of an image, a standard pt-PT transcription prompt, and the ground-truth transcription. The dataset is licensed under CC BY-NC 4.0, for non-commercial use only, and proper attribution is required.




