HePU: Hebrew Paleography Understanding Dataset For Semantic Script Segmentation
收藏资源简介:
We introduce HePU, a novel dataset for the semantic segmentation of historical Hebrew script styles. Derived from Sfardata, the codicological database of the Hebrew Palaeography Project, HePU provides high-quality, pixel-level annotations for Hebrew manuscript analysis. The dataset comprises 193 manuscript page images, each annotated with precise semantic segmentation labels. Every page corresponds to one of six major Hebrew script types: Ashkenazi, Byzantine, Italian, Oriental, Sefardic, and Yemenite. By ensuring that each page contains only a single script type, HePU provides unambiguous labels and enables models to learn script-specific visual features effectively. HePU is designed to support research in Hebrew paleography, facilitating tasks such as script classification, layout analysis, and historical manuscript preservation.



