Dataset for: Towards Autonomous Epigraphy: End-to-End Transliteration of Cuneiform Tablets from 2D Photographs using Vision-Language Models
收藏资源简介:
This repository contains the dataset splits used in the study "Towards Autonomous Epigraphy: End-to-End Transliteration of Cuneiform Tablets from 2D Photographs using Vision-Language Models." Overview: The dataset consists of a quality-filtered corpus of 6,116 pairs of cuneiform tablet identifiers and their corresponding texts. It is structured into pre-defined train, validation, and test splits to ensure reproducibility of the Vision-Language Model (VLM) training and evaluation described in the paper. Format: The data is provided as a HuggingFace Dataset (.arrow format) encompassing the necessary splits. Data Sources: The original cuneiform tablet photographs and ASCII Transliteration Format (ATF) transliterations that map to these splits are provided by and available through the Electronic Babylonian Library (eBL) portal (https://www.ebl.lmu.de).



