遇见数据集

amalia-llm/DocVQA-PT

收藏
Hugging Face2026-07-06 更新2026-07-22 收录
官方服务:

资源简介:

--- language: - pt license: cc-by-nc-4.0 source_datasets: - lmms-lab/DocVQA pretty_name: DocVQA-PT task_categories: - image-text-to-text configs: - config_name: default data_files: - split: validation path: data/validation-* dataset_info: features: - name: questionId dtype: string - name: question dtype: string - name: question_types list: string - name: image dtype: image - name: docId dtype: int64 - name: ucsf_document_id dtype: string - name: ucsf_document_page_no dtype: string - name: answers list: string - name: data_split dtype: string splits: - name: validation num_bytes: 3579446465.63 num_examples: 5349 download_size: 1056534286 dataset_size: 3579446465.63 --- <div align="center"> <img width="400px" src="https://github.com/AMALIA-LLM/amalia-llm.github.io/blob/main/source/_static/logo/logo-color-black.png?raw=true"> [![AMALIA](https://img.shields.io/badge/AMALIA-Website-1A1A1A?labelColor=FDE000&style=for-the-badge)](https://amaliallm.pt/) [![Paper](https://img.shields.io/badge/%20-Paper-1A1A1A?labelColor=FDE000&logo=googlescholar&logoColor=black&style=for-the-badge)](https://arxiv.org/abs/2606.19100) [![Data](https://img.shields.io/badge/%20-Data-1A1A1A?labelColor=FDE000&logo=huggingface&logoColor=black&style=for-the-badge)](https://huggingface.co/collections/amalia-llm/amalia-vl-eval) [![Main Repo](https://img.shields.io/badge/%20-Main%20Repo-1A1A1A?labelColor=FDE000&logo=github&logoColor=black&style=for-the-badge)](https://github.com/AMALIA-LLM/AMALIA) [![Eval Repo](https://img.shields.io/badge/%20-Eval%20Repo-1A1A1A?labelColor=FDE000&logo=github&logoColor=black&style=for-the-badge)](https://github.com/AMALIA-LLM/amalia-vl-eval) </div> # DocVQA-PT European Portuguese (pt-PT) machine translation of **DocVQA**, a visual question answering dataset over scanned document images. Translated from the original English `validation` split (`DocVQA` subset) using **gemini-3.1-pro**. **Original Dataset:** [https://huggingface.co/datasets/lmms-lab/DocVQA](https://huggingface.co/datasets/lmms-lab/DocVQA) (`DocVQA` subset) **Note:** This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the [AMALIA project](https://github.com/AMALIA-LLM/AMALIA) and is included in [amalia-vl-eval](https://github.com/AMALIA-LLM/amalia-vl-eval), a comprehensive benchmark suite for evaluating large vision-language models on European Portuguese. --- ## Licensing This dataset is **dual-licensed**: - **Translated content** — the questions, answers, and any other text translated to European Portuguese with **gemini-3.1-pro** is released under [`cc-by-nc-4.0`](https://creativecommons.org/licenses/by-nc/4.0/) (Attribution–NonCommercial). It may be used for non-commercial purposes with attribution. - **Images and untranslated columns** — all images and any columns that were not translated retain the original **DocVQA** license (refer to the [original dataset](https://huggingface.co/datasets/lmms-lab/DocVQA)). When using this dataset you must comply with **both** licenses. The NonCommercial terms of the translated content do not override any stricter conditions of the original benchmark's license, and vice versa. --- ## Citation If you use this dataset or AMALIA-VL in your work, please cite: ```bibtex @article{gloria2026amalia, title={AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model}, author={Gl{\'o}ria-Silva, Diogo and Cardeira, Jo{\~a}o and da Luz, Manuel Letras and Simpl{\'\i}cio, Afonso and Vinagre, Gon{\c{c}}alo and Tavares, Diogo and Ferreira, Rafael and Calvo, In{\^e}s and Vieira, In{\^e}s and Semedo, David and others}, journal={arXiv preprint arXiv:2606.19100}, year={2026} } ```

European Portuguese (pt-PT) machine translation of DocVQA, a visual question answering dataset over scanned document images. Translated from the original English validation split (DocVQA subset) using gemini-3.1-pro. This dataset is part of the AMALIA project and included in amalia-vl-eval, a benchmark suite for evaluating large vision-language models on European Portuguese.

提供机构:
amalia-llm
搜集汇总
数据集介绍
amalia-llm/DocVQA-PT 数据集图片
构建方式
DocVQA-PT数据集基于广受认可的视觉问答数据集DocVQA构建,专注于扫描文档图像的视觉问答任务。该数据集通过先进的大语言模型gemini-3.1-pro,将DocVQA原始英文验证集(DocVQA子集)中的问题、答案及所有文本内容,精准翻译为欧洲葡萄牙语(pt-PT)。翻译过程严格保留了原始数据集的图像、文档标识等非文本字段,确保了多模态信息的完整性。最终形成一个包含5349个样本的高质量验证集,服务于葡萄牙语场景下的文档视觉问答研究。
特点
DocVQA-PT数据集的显著特点在于其语言针对性与评估实用性。作为AMALIA项目的重要组成部分,它专门面向欧洲葡萄牙语环境,填补了该语言在文档视觉问答领域的空白。数据集完全继承DocVQA的扫描文档图像多样性,涵盖表格、发票、信件等多种真实文档类型,问题设计涉及文本检索、数值推理、语义理解等多个层次。此外,其双许可机制(CC-BY-NC 4.0与原始许可)兼顾了翻译内容与原始内容的版权保护,为研究社区提供了清晰的使用边界。
使用方法
DocVQA-PT数据集的使用方法直观且灵活。研究人员可直接通过HuggingFace平台加载数据,利用标准的视觉问答评测流程评估模型性能。该数据集与AMALIA项目的评估套件amalia-vl-eval无缝集成,支持自动化的多模型对比评测。使用时需注意遵守双许可条款,非商业应用需标注翻译内容来源。对于评估大型视觉语言模型在欧洲葡萄牙语场景下的文档理解能力,本数据集提供了标准化的验证基准,支持零样本评测与微调实验的多种范式。
背景与挑战
背景概述
DocVQA-PT数据集由葡萄牙AMALIA项目团队于2026年创建,旨在将广泛使用的DocVQA文档视觉问答基准扩展至欧洲葡萄牙语。该数据集基于原始DocVQA验证集,采用Gemini 3.1 Pro进行机器翻译,生成共计5349个问答样本,涵盖扫描文档图像的理解任务。作为AMALIA-VL评估套件的核心组成部分,DocVQA-PT致力于推动多语言视觉语言模型在低资源语言场景下的性能评估,为葡萄牙语自然语言处理与文档智能分析交叉领域提供了标准化测试平台。该数据集的发布填补了现有文档视觉问答基准在葡语语种上的空白,对促进多模态人工智能系统的语言包容性具有显著学术价值。
当前挑战
DocVQA-PT首先面临文档视觉问答领域的两大核心挑战:其一是复杂文档图像中文本布局与排版多样性导致的视觉理解困难,包括表格、多栏排版、手写注释等要素的准确解析;其二是开放式问题对模型语义推理与细粒度信息定位能力的高要求。在构建过程中,机器翻译固有的局限性构成主要挑战:Gemini模型可能引入学术或技术术语的葡萄牙语翻译歧义,以及合成样本中问答逻辑一致性的潜在偏差。此外,原始DocVQA仅保留验证集子集进行翻译,样本规模有限,可能影响模型在不同文档类型上的泛化能力评估。双许可证结构亦要求使用者需同时遵守CC-BY-NC-4.0与原始基准许可证的复合条款。
常用场景
经典使用场景
DocVQA-PT数据集作为视觉问答领域中的标杆性资源,专门面向欧洲葡萄牙语场景。其经典用途是在扫描文档图像上执行问答任务,要求模型同时理解图像中的文字布局与语义内容,并基于葡萄牙语问题给出精确答案。该数据集将原始DocVQA的验证集通过机器翻译转化为葡萄牙语版本,保留了图像与文档结构的多样性,成为评估多模态模型在多语言、低资源环境下的文本理解与推理能力的理想测试平台。
衍生相关工作
围绕DocVQA-PT衍生了一系列推动多语言多模态发展的经典工作。其中,AMALIA项目将其纳入amalia-vl-eval基准套件,系统评估了多个视觉语言模型在葡萄牙语文档问答中的表现,并发布了对应的开源模型AMALIA-VL。此外,研究者基于该数据集探索了机器翻译质量对问答性能的影响,以及如何利用数据增强和跨语言微调策略来弥合语言偏差,进而催生了面向低资源语言的视觉问答评估框架和多样化翻译校正方法。
数据集最近研究
最新研究方向
DocVQA-PT的推出标志着视觉问答领域向低资源语言纵深拓展的前沿动向,其以欧洲葡萄牙语对文档级视觉问答基准DocVQA进行全面机器翻译,旨在填补非英语场景下文档理解评估的空白。结合AMALIA项目构建的amalia-vl-eval综合评测套件,该数据集聚焦于多模态大语言模型在葡语文档图像上的鲁棒性与语义对齐能力,呼应了全球范围内对语言多样性AI评估的热切关注。这一科研实践不仅为葡语社区提供了关键基准,更激励学术界重新审视现有视觉语言模型在跨语言泛化中的偏见与局限,推动构建更为公平、包容的多模态理解系统。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务