OCR model for Pracalit for Sanskrit and Newar MSS 16th to 19th C., Ground Truth
收藏资源简介:
Ground truth data (png and xml files) for a an OCR model. Will be continually updated. Originally trained on Transkribus with a PyLaia model created from ground truth data based on transcripts into Pracalit Unicode of four Nepalese manuscripts. The manuscripts used to create this model are Staatsbibliothek zu Berlin's Hitopadeśa (MIK I 4851) (mixed Newar and Sanskrit dating to 1561) and Vetālapañcaviṃśati (HS. Or. 6414) (Newar dating to 1675) as well as Cambridge Digital Library's Avalokiteśvaraguṇakāraṇḍavyūha (MS Add. 1322) (Sanskrit, 18th century) and the Royal Asiatic Society Online Collection's Madhyamasvayaṃbhūpurāṇa (RAS Hodgson MS 23) (Newar and Sanskrit dating to c. 1800). The training was done on 441 pages and validation on 242 pages. This model does not recognise spacing, except for large gaps (i.e. for pictures or string holes). Newar word divider markers may not be represented or may be transcribed as virama. In general, the model is made for MSS with scriptio continua and will transcribe into scriptio continua into Pracalit Unicode. Transcription was performed by Dr Alexander O'Neill (SOAS University of London). Transcription of the Vetālapañcaviṃśati (HS. Or. 6414) and Madhyamasvayaṃbhūpurāṇa (RAS Hodgson MS 23) was aided by unpublished materials provided by Dr Felix Otter (Philipps-Universität Marburg), as well as the published transcription in Shakya, Min Bahadur, and Shanta Harsha Bajracharya, eds. "Svayambhū Purāṇa." Lalitpur: Nagarjuna Institute of Exact Methods, 2001. The transcription of Avalokiteśvaraguṇakāraṇḍavyūha (MS Add. 1322) was aided by the transcription provided by the Digital Sanskrit Buddhist Canon Project based on Lokesh Chandra, "Guṇakāraṇḍavyūhasūtram," New Delhi: International Academy of Indian Culture, 1999.
本数据集为光学字符识别(Optical Character Recognition, OCR)模型的真实标注数据集,包含PNG与XML格式文件,且将持续更新。该模型最初依托四份尼泊尔手稿的转录文本,通过Transkribus平台,基于本真实标注数据构建PyLaia模型,并将转录结果转换为Pracalit Unicode编码。 用于构建该模型的手稿包括:柏林国家图书馆(Staatsbibliothek zu Berlin)藏《五卷书》(*Hitopadeśa*,编号MIK I 4851),为混合尼瓦尔语与梵语文本,年代可追溯至1561年;《鬼话二十五则》(*Vetālapañcaviṃśati*,编号HS. Or. 6414),为尼瓦尔语文本,年代为1675年;剑桥数字图书馆(Cambridge Digital Library)藏《观自在功德藏经》(*Avalokiteśvaraguṇakāraṇḍavyūha*,编号MS Add. 1322),为梵语文本,创作于18世纪;以及皇家亚洲学会在线馆藏(Royal Asiatic Society Online Collection)藏《中自生往世书》(*Madhyamasvayaṃbhūpurāṇa*,编号RAS Hodgson MS 23),为混合尼瓦尔语与梵语文本,年代约为1800年。 本次模型训练共使用441页文本,验证集包含242页。该模型无法识别常规字符间距,仅可识别大幅空白(如用于配图或字符串留白的间隙)。尼瓦尔语单词分隔符可能无法被正确标注,或被转录为维尔玛(virama)符号。总体而言,本模型专为采用连续书写体(scriptio continua)的手稿设计,转录结果将以连续书写格式输出为Pracalit Unicode编码。 本次转录工作由伦敦大学亚非学院(SOAS University of London)的亚历山大·奥尼尔博士(Dr Alexander O'Neill)完成。其中,《鬼话二十五则》(HS. Or. 6414)与《中自生往世书》(RAS Hodgson MS 23)的转录工作得到了马尔堡菲利普斯大学(Philipps-Universität Marburg)费利克斯·奥特博士(Dr Felix Otter)提供的未刊材料,以及由沙克亚·明·巴哈杜尔与香塔·哈尔沙·巴特拉查里亚编纂的*Svayambhū Purāṇa*,拉利特普尔:那加朱纳精确方法研究所,2001年。《观自在功德藏经》(MS Add. 1322)的转录则得益于梵语佛教数字藏经项目(Digital Sanskrit Buddhist Canon Project)提供的转录文本,该转录基于洛凯什·钱德拉所著*Guṇakāraṇḍavyūhasūtram*,新德里:印度文化国际学院,1999年。



