遇见数据集

NABaybay Dataset

收藏
Zenodo2025-08-10 更新2026-05-26 收录
官方服务:

资源简介:

The Novel Assembled Baybayin (NABaybay) Dataset is a curated collection of Baybayin script data designed for computer vision applications and related research. It contains 61,200 Baybayin character images, 1,000 Baybayin word images, and 110 Baybayin block-level images, each prepared in formats suitable for various machine learning and deep learning workflows. The character images are available in four formats: raw (.jpg), grayscale (.jpg), feature vector (.csv), and feature vector (.mat). Each Baybayin character class contains 3,600 images, processed into binary, center-aligned, and 28×28-pixel data. The word image set consists of 1,000 raw (.png) files, each named according to its Latin equivalent. The block image set includes over 100 raw (.png) images, comprising 55 handwritten and 35 typewritten Baybayin texts, some of which also include accompanying Latin text. The source code and scripts used to generate this dataset are available in our GitHub repository: https://github.com/rbp0803/NABaybay_Dataset This dataset is part of an ongoing initiative to preserve and reintroduce the Baybayin script in the digital age. It can be applied to a wide range of Baybayin-related programs, studies, and research projects.

提供机构:
Zenodo
创建时间:
2025-08-10
二维码
社区交流群
二维码
科研交流群
商业服务