Jibo-labeled classical hiragana image dataset
收藏资源简介:
JIBO-LABELED CLASSICAL HIRAGANA IMAGE DATASET This dataset is intended to be used to train an optical character recognition (OCR) model to recognize hiragana 平仮名 in classical Japanese manuscripts and woodblock printed books. Hiragana are Japanese phonetic symbols derived from cursivized forms of Chinese characters. Among them, so-called hentaigana 変体仮名 (lit. variant kana) are forms of hiragana that have been more-or-less obsolete since 1900. Decoding hentaigana is a non-trivial task that is essential for text extraction from images of classical Japanese manuscripts, i.e. those inscribed before 1900. In order to train a deep-learning model that can recognize individual jibo variants, we modified the hiragana section of an existing dataset called the Kotenseki kuzushiji dataset (Version 2, 2019, http://codh.rois.ac.jp/char-shape/, consisting of 44 items, 6,151 scans, 4,328 character types, and 1,086,326 individual characters) and refined it with jibo labels. Rather than adding the jibo labels by hand, we used the Kokugoken hentaigana jikei database (consisting of 7 items) to train a small model that automatically added jibo labels, which we then corrected by human inspection. Note: as of this writing (20 April 2026), the CODH website is offline, so we provide a working link for the former dataset through the internet archive: https://web.archive.org/web/20250827181728/https://codh.rois.ac.jp/char-shape/. The Kotenseki kuzushiji dataset provides coordinate information for another dataset consisting of scans of classical Japanese manuscripts, called the Kotenseki dataset. Both datasets were jointly developed by the National Institute of Japanese Literature (NIJL) and the Center for Open Data in the Humanities (CODH). Regarding the contents of the Kotenseki dataset, please see the following page: https://codh.rois.ac.jp/char-shape/book/ (accessed 27 August 2025). Note: as of this writing (20 April 2026), the CODH website is offline, so we provide a working link through the internet archive: https://web.archive.org/web/20250827181728/https://codh.rois.ac.jp/char-shape/book/. The current Jibo-labeled classical hiragana image dataset contains the hiragana portion of the Kotenseki kuzushiji dataset with jibo labels. It was derived from the Kotenseki kuzushiji dataset, with bootstrapping of the labeling from the Kokugoken hentaigana jikei database. We made the best effort to correct auto-labeled data, but potential mistakes remain in the current version. We are releasing this dataset in hopes that other will be able to use it to train other models or to add jibo-decoding functionalities to their existing classical Japanese OCR models. For more details, please see the file README.md. ACKNOWLEDGEMENTS We express our deepest appreciation to: • the National Institute for Japanese Literature (Kokubungaku kenkyu shiryokan 国文学研究資料館) for creating, tagging, and sharing the images from their collections; • the Center for Open Data in the Humanities (Jinbungaku oopun deeta kyodoriyo sentaa 人文学オープンデータ共同利用センター) for processing and releasing the original dataset; • the National Institute for Japanese Language and Linguistics (Kokuritsu kokugo kenkyusho 国立国語研究所) for sharing their valuable data online. We thank Amazon Web Services (AWS) and the eScience Institute of the University of Washington, Seattle, for providing cloud computing credits. We thank the Simpson Center for the Humanities, University of Washington, Seattle for Digital Humanities Summer Fellowships (Summer 2025). ATTRIBUTION If you reuse this dataset, the following citation pattern is recommended: "Michael R. Zeng, Herman Chau, and Paul S. Atkins (2026). Jibo-labeled classical hiragana image dataset. doi:10.5281/zenodo.18765466." LICENSE AND DISCLAIMERS This dataset is bound by the terms of the underlying Nihon kotenseki kuzushiji deetasetto, which was publicly released under CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/). The underlying dataset is provided by ROIS-DS Center for Open Data in the Humanities (CODH) (https://web.archive.org/web/20250827113242/http://codh.rois.ac.jp/). Here is the recommended citation for the original dataset: "Nihon kotenseki kuzushiji deetasetto (from the collection of NIJL and others, processed by CODH) 『日本古典籍くずし字データセット』 (国文研ほか所蔵/CODH加工) doi:10.20676/00000340" Corrections and minor changes were made to the original dataset, in addition to the jibo tags. The creators of the original dataset do not necessarily endorse this derivative. No warranties are given regarding this dataset. We cannot provide technical support. If you find the dataset useful, please let us know! We would love to hear from you.
JIBO标注古典平假名图像数据集(JIBO-LABELED CLASSICAL HIRAGANA IMAGE DATASET) 本数据集旨在用于训练光学字符识别(Optical Character Recognition,OCR)模型,以识别日本古典手稿与木刻印刷书籍中的平假名。平假名是源自汉字草书形态的日语表音符号。其中,所谓的变体假名(hentaigana)即自1900年起基本废弃的平假名书写形式。对变体假名的解码是一项极具挑战性的任务,也是从1900年之前的日本古典手稿图像中提取文本的核心必要步骤。 为训练能够识别各类字形变体的深度学习模型,我们对现有名为古典籍崩字数据集(Kotenseki kuzushiji dataset)(2019年第2版,http://codh.rois.ac.jp/char-shape/,包含44项内容、6151幅扫描图、4328种字符类型以及1086326个独立字符)的平假名部分进行了修改,并添加了JIBO标注。我们并未手动添加JIBO标注,而是借助国文研变体假名字形数据库(Kokugoken hentaigana jikei database)(包含7项内容)训练了一个小型模型,以自动完成JIBO标注,随后通过人工审核对标注结果进行了修正。注:截至本文撰写时(2026年4月20日),CODH官网已下线,因此我们通过互联网档案馆提供该原始数据集的可用链接:https://web.archive.org/web/20250827181728/https://codh.rois.ac.jp/char-shape/。 古典籍崩字数据集(Kotenseki kuzushiji dataset)提供了另一项数据集的坐标信息,该数据集名为古典籍数据集,包含日本古典手稿的扫描图像。这两项数据集均由日本国文厅(National Institute of Japanese Literature, NIJL)与人文科学开放数据中心(Center for Open Data in the Humanities, CODH)联合开发。关于古典籍数据集的详细内容,请参阅以下页面:https://codh.rois.ac.jp/char-shape/book/(2025年8月27日访问)。注:截至本文撰写时(2026年4月20日),CODH官网已下线,因此我们通过互联网档案馆提供该页面的可用链接:https://web.archive.org/web/20250827181728/https://codh.rois.ac.jp/char-shape/book/。 当前的JIBO标注古典平假名图像数据集包含古典籍崩字数据集(Kotenseki kuzushiji dataset)中的平假名部分,并添加了JIBO标注。本数据集源自古典籍崩字数据集(Kotenseki kuzushiji dataset),标注工作通过国文研变体假名字形数据库(Kokugoken hentaigana jikei database)进行引导式标注。我们已尽最大努力修正自动标注的数据,但当前版本中仍可能存在错误。 我们发布本数据集,希望其他研究者能够利用它训练新的模型,或为现有日本古典OCR模型添加变体假名解码功能。 更多细节请参阅README.md文件。 致谢 我们谨向以下机构致以最诚挚的谢意: • 日本国文厅(National Institute of Japanese Literature):为其馆藏图像的创建、标注与共享提供支持; • 人文科学开放数据中心(Center for Open Data in the Humanities):对原始数据集进行处理并发布; • 国立国语研究所(Kokuritsu kokugo kenkyusho):在线共享其珍贵的数据集。 我们感谢亚马逊云服务(Amazon Web Services, AWS)与华盛顿大学西雅图分校电子科学研究所提供云计算额度。我们感谢华盛顿大学西雅图分校辛普森人文中心为我们提供的数字人文夏季研究员奖学金(2025年夏季)。 引用规范 若您复用本数据集,推荐采用以下引用格式: "Michael R. Zeng, Herman Chau, 和 Paul S. Atkins (2026). JIBO标注古典平假名图像数据集. doi:10.5281/zenodo.18765466." 许可与免责声明 本数据集受日本古典籍崩字数据集(Nihon kotenseki kuzushiji deetasetto)的许可条款约束,该原始数据集以CC BY-SA 4.0(https://creativecommons.org/licenses/by-sa/4.0/)协议公开发布。原始数据集由ROIS-DS人文科学开放数据中心(CODH)提供(https://web.archive.org/web/20250827113242/http://codh.rois.ac.jp/)。原始数据集的推荐引用格式如下: "日本古典籍くずし字データセット(国文研ほか所蔵/CODH加工) doi:10.20676/00000340" 除添加JIBO标注外,我们还对原始数据集进行了修正与少量修改。原始数据集的创作者未必认可本衍生数据集。 本数据集不提供任何担保,我们无法提供技术支持。 若您认为本数据集对您有所帮助,请务必告知我们!我们期待收到您的反馈。



