Cretogram Dataset of Isolated Characters from Archaic and Classical Crete
收藏资源简介:
Cretogram is a dataset of 11,611 isolated grayscale character images from 187 Cretan inscription sources, organized into 25 classes and normalized to 92 × 92 pixels. It supports research on character recognition in ancient Greek epigraphy. The collection includes Eteocretan-language inscriptions written in local Cretan alphabets. Left-facing and right-facing variants of the same letter share a class. The images were prepared through manual background-noise removal, glyph extraction, normalization and character annotation. The deposit includes PNG images, NumPy arrays, source-provenance metadata, predefined partitions, Python code and saved benchmark predictions. The benchmark comprises 11,583 images in 22 classes, with five inscription-disjoint folds and a separate compact-CNN control allowing inscription overlap. Baseline results are provided for HOG–SVM, a compact CNN and pretrained ResNet-18. Current annotations include one documented correction; the labels used in the reported experiments are retained separately. Preparation methods, evaluation protocols and annotation history are described in the accompanying README and supplementary documentation. Original contributions are licensed under Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0), subject to the source-material rights specified in the package.



