amalia-llm/DocVQA-PT
收藏资源简介:
--- language: - pt license: cc-by-nc-4.0 source_datasets: - lmms-lab/DocVQA pretty_name: DocVQA-PT task_categories: - image-text-to-text configs: - config_name: default data_files: - split: validation path: data/validation-* dataset_info: features: - name: questionId dtype: string - name: question dtype: string - name: question_types list: string - name: image dtype: image - name: docId dtype: int64 - name: ucsf_document_id dtype: string - name: ucsf_document_page_no dtype: string - name: answers list: string - name: data_split dtype: string splits: - name: validation num_bytes: 3579446465.63 num_examples: 5349 download_size: 1056534286 dataset_size: 3579446465.63 --- <div align="center"> <img width="400px" src="https://github.com/AMALIA-LLM/amalia-llm.github.io/blob/main/source/_static/logo/logo-color-black.png?raw=true"> [](https://amaliallm.pt/) [](https://arxiv.org/abs/2606.19100) [](https://huggingface.co/collections/amalia-llm/amalia-vl-eval) [](https://github.com/AMALIA-LLM/AMALIA) [](https://github.com/AMALIA-LLM/amalia-vl-eval) </div> # DocVQA-PT European Portuguese (pt-PT) machine translation of **DocVQA**, a visual question answering dataset over scanned document images. Translated from the original English `validation` split (`DocVQA` subset) using **gemini-3.1-pro**. **Original Dataset:** [https://huggingface.co/datasets/lmms-lab/DocVQA](https://huggingface.co/datasets/lmms-lab/DocVQA) (`DocVQA` subset) **Note:** This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the [AMALIA project](https://github.com/AMALIA-LLM/AMALIA) and is included in [amalia-vl-eval](https://github.com/AMALIA-LLM/amalia-vl-eval), a comprehensive benchmark suite for evaluating large vision-language models on European Portuguese. --- ## Licensing This dataset is **dual-licensed**: - **Translated content** — the questions, answers, and any other text translated to European Portuguese with **gemini-3.1-pro** is released under [`cc-by-nc-4.0`](https://creativecommons.org/licenses/by-nc/4.0/) (Attribution–NonCommercial). It may be used for non-commercial purposes with attribution. - **Images and untranslated columns** — all images and any columns that were not translated retain the original **DocVQA** license (refer to the [original dataset](https://huggingface.co/datasets/lmms-lab/DocVQA)). When using this dataset you must comply with **both** licenses. The NonCommercial terms of the translated content do not override any stricter conditions of the original benchmark's license, and vice versa. --- ## Citation If you use this dataset or AMALIA-VL in your work, please cite: ```bibtex @article{gloria2026amalia, title={AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model}, author={Gl{\'o}ria-Silva, Diogo and Cardeira, Jo{\~a}o and da Luz, Manuel Letras and Simpl{\'\i}cio, Afonso and Vinagre, Gon{\c{c}}alo and Tavares, Diogo and Ferreira, Rafael and Calvo, In{\^e}s and Vieira, In{\^e}s and Semedo, David and others}, journal={arXiv preprint arXiv:2606.19100}, year={2026} } ```
European Portuguese (pt-PT) machine translation of DocVQA, a visual question answering dataset over scanned document images. Translated from the original English validation split (DocVQA subset) using gemini-3.1-pro. This dataset is part of the AMALIA project and included in amalia-vl-eval, a benchmark suite for evaluating large vision-language models on European Portuguese.




