遇见数据集

openpecha/uchen_ume_classification_dataset

收藏
Hugging Face2026-05-26 更新2026-06-14 收录
官方服务:

资源简介:

Uchen-Ume分类基准是一个二进制图像分类数据集,专门用于区分藏文的两种基本书写类别:Uchen(དབུ་ཅན།,有头字,具有水平顶部笔画)和Ume(དབུ་མེད།,无头字,无顶部笔画)。数据集包含来自佛教数字资源中心(BDRC)的原始未处理手稿扫描图像,覆盖了多样化的条件,如老化纸张、现代重印、不同墨水密度、扫描设备差异以及多个世纪的手稿生产。数据集总共有10,961个示例,分为训练集(9,110个示例)、验证集(1,000个示例)和测试集(851个示例),并按照工作级别进行分层分割以防止数据泄漏。图像以原始分辨率和宽高比存储(通常为5:1的横向pecha格式),未进行任何预处理(如调整大小、裁剪或归一化),以最大化数据集的复用性。注释过程基于藏文书写类型的多年分类学,通过结构化流程完成,并经过质量控制。该数据集适用于图像分类、古文字学、藏文手稿研究等领域,旨在促进藏文脚本自动分类模型的发展。

The Uchen-Ume Classification Benchmark is a binary image classification dataset for distinguishing two fundamental categories of Tibetan script: Uchen (དབུ་ཅན།, headed script with a horizontal top stroke) and Ume (དབུ་མེད།, headless script without a top stroke). All images are raw, unprocessed manuscript scans from the Buddhist Digital Resource Center (BDRC), encompassing a wide range of conditions such as aged paper, modern reprints, varying ink densities, different scanning equipment, and multiple centuries of manuscript production. The dataset contains a total of 10,961 examples, split into training (9,110 examples), validation (1,000 examples), and test (851 examples) sets, with splits stratified by class and partitioned at the work level to prevent data leakage. Images are stored in their original resolution and aspect ratio (typically 5:1 landscape pecha format) without any preprocessing (e.g., resizing, cropping, or normalization) to maximize reusability across different experimental setups. The annotation process is based on a multi-year typology of Tibetan scripts, developed through a structured workflow with quality control. This dataset is designed for image classification, paleography, and Tibetan Buddhist manuscript research, aiming to advance automatic script classification models for Tibetan texts.

提供机构:
openpecha
二维码
社区交流群
二维码
科研交流群
商业服务