Pelerinage: Dataset
收藏资源简介:
The "<em>Pelerinage"</em> dataset contains a fine-grained edition of excerpts from 20 medieval manuscripts of Guillaume de Digulleville's <em>Pelerinage de vie humaine.</em> Files in this directory were created as part of the <em>ECMEN Ecritures médiévales et outils numériques</em> research project, funded by the City of Paris in the <em>Emergence(s)</em> framework. If you use any of the following files, please quote: Stutzmann, Dominique. « Les 'manuscrits datés', base de données sur l’écriture ». In Catalogazione, storia della scrittura, storia del libro. I Manoscritti datati d’Italia vent’anni dopo, ed. Teresa De Robertis and Nicoletta Giovè Marchioli, Firenze: SISMEL - Edizioni del Galluzzo, 2017, p. 155-207. <pre>@incollection{stutzmann_les_2017, address = {Firenze}, title = {Les « manuscrits datés », base de données sur l’écriture}, language = {fre}, booktitle = {Catalogazione, storia della scrittura, storia del libro. {I} {Manoscritti} datati d’{Italia} vent’anni dopo}, publisher = {SISMEL - Edizioni del Galluzzo}, author = {Stutzmann, Dominique}, editor = {De Robertis, Teresa and Giovè Marchioli, Nicoletta}, year = {2017}, pages = {155--207} } </pre> <strong>SOURCE DESCRIPTION</strong> The source of the files are imitative transcriptions of selected passages by Géraldine Veysseyre (ORCID 0000-0002-3737-2137) based on 20 medieval manuscripts containing the "Pelerinage de vie humaine" of Guillaume de Digulleville, as produced during the OPVS research project (the acronym stands for « Old Pious Vernacular Successes » in English and « Œuvres Pieuses Vernaculaires à Succès » in French), funded by the ERC under the grant agreement n° 263274). The transcriptions were edited and enhanced in TEI format by Dominique Stutzmann (ORCID 0000-0003-3705-5825) and Floriana Ceresato (IRHT-CNRS), as part of the ORIFLAMMS and the ECMEN (<em>Ecritures médiévales et outils numériques) </em>research project with following features: semantic encoding (<persName/>, <placeName/>, <title/>, <num/>) <choice/>, <abbr/>, <expan/> etc. and entities to describe abbreviations minimal description of manuscripts Links to images and coordinates to pixels of the image for each line were added by Floriana Ceresato and Alexandre Gaudin (IRHT-CNRS). Files in ALTO format were produced in 2021 by Jean-Baptiste Camps (ORCID 0000-0003-0385-7037) and Chahan Vidal-Gorène (ORCID 0000-0003-1567-6508) as part of a joint research presented at the conference DH2022 (Tokyo, 25-29 July 2022), cf. (abstract) Camps, Jean-Baptiste, Chahan Vidal-Gorène, Dominique Stutzmann, Marguerite Vernet, and Ariane Pinche. « Data Diversity in Handwritten Text Recognition: Challenge or Opportunity? » In <em>Digital Humanities 2022. Conference Abstracts (The University of Tokyo, Japan, 25-29 July 2022)</em>, published by DH2022 Local Organizing Committee, 160‑65. Tokyo, 2022. https://dh2022.dhii.asia/dh2022bookofabsts.pdf#page=162 and https://dh2022.dhii.asia/abstracts/files/CAMPS_Jean_Baptiste_Data_Diversity_in_handwritten_text_recog.html. The dataset comprises the following folders: <strong>/texts/</strong> : main TEI file containing all textual, semantic and graphic information. <strong>/img/</strong> : 49 images on which the transcriptions are based and to which they are aligned at line level through the <facsimile/> element in the TEI file. <strong>/alto/</strong> : one file per image in Alto format in several flavours according to the needs of the users during the experiments of the above mentioned paper. The text is flat, line by line, either with expansion or with abbreviations, and either with standardisation or according to the original encoding. Given the current implementation of Kraken/eScriptorium, coordinates may be indicated as "x1,y1 x2,y2..." in "/alto/without-norm-coord-commas/" but as "x1 y1 x2 y2..." in the other folders. <strong>/img-masks/ </strong>: smaller images with a view of all coordinates of lines as present in the TEI > facsimile elements, including lines which were not selected for the transcription.
<em>《朝圣》(Pelerinage)</em>数据集包含纪尧姆·德·迪居勒维尔(Guillaume de Digulleville)所著《人生朝圣》(Pelerinage de vie humaine)20部中世纪手稿节选的细粒度校勘版本。本目录下的文件由<em>中世纪书写与数字工具(ECMEN, Ecritures médiévales et outils numériques)</em>研究项目制作,该项目由巴黎市在"Emergence(s)"资助框架下立项资助。若使用本数据集下的任意文件,请引用如下文献:Stutzmann, Dominique. « Les 'manuscrits datés', base de données sur l’écriture ». In Catalogazione, storia della scrittura, storia del libro. I Manoscritti datati d’Italia vent’anni dopo, ed. Teresa De Robertis and Nicoletta Giovè Marchioli, Firenze: SISMEL - Edizioni del Galluzzo, 2017, p. 155-207. <pre>@incollection{stutzmann_les_2017, address = {Firenze}, title = {Les « manuscrits datés », base de données sur l’écriture}, language = {fre}, booktitle = {Catalogazione, storia della scrittura, storia del libro. {I} {Manoscritti} datati d’{Italia} vent’anni dopo}, publisher = {SISMEL - Edizioni del Galluzzo}, author = {Stutzmann, Dominique}, editor = {De Robertis, Teresa and Giovè Marchioli, Nicoletta}, year = {2017}, pages = {155--207} }</pre> <strong>来源说明</strong> 本数据集文件的文本源自热拉尔丁·韦瑟雷尔(Géraldine Veysseyre,ORCID 0000-0002-3737-2137)基于20部收录纪尧姆·德·迪居勒维尔《人生朝圣》的中世纪手稿制作的精选段落摹写转录本,该转录工作由OPVS研究项目(全称为英文*Old Pious Vernacular Successes*、法文*Œuvres Pieuses Vernaculaires à Succès*,即“古通俗圣典传世之作”)完成,该项目由欧洲研究委员会(ERC)根据编号为263274的资助协议立项资助。上述转录本由多米尼克·施图茨曼(Dominique Stutzmann,ORCID 0000-0003-3705-5825)与弗洛里亚娜·切雷萨托(Floriana Ceresato,IRHT-CNRS)在ORIFLAMMS及<em>中世纪书写与数字工具(ECMEN, Ecritures médiévales et outils numériques)</em>研究项目框架下进行TEI格式的编辑与优化,具体包含以下功能:语义编码(<persName/>、<placeName/>、<title/>、<num/>)、<choice/>、<abbr/>、<expan/>等标记,以及用于描述缩写与手稿极简信息的实体标注。弗洛里亚娜·切雷萨托与亚历山大·戈丹(Alexandre Gaudin,IRHT-CNRS)为每一行文本添加了对应图像的链接及像素坐标信息。ALTO格式文件由让-巴蒂斯特·坎普斯(Jean-Baptiste Camps,ORCID 0000-0003-0385-7037)与沙汉·维达尔-戈雷内(Chahan Vidal-Gorène,ORCID 0000-0003-1567-6508)于2021年制作,相关研究成果发表于2022年7月25日至29日在日本东京举办的DH2022数字人文国际会议,具体可参考如下摘要:Camps, Jean-Baptiste, Chahan Vidal-Gorène, Dominique Stutzmann, Marguerite Vernet, and Ariane Pinche. « Data Diversity in Handwritten Text Recognition: Challenge or Opportunity? » In <em>Digital Humanities 2022. Conference Abstracts (The University of Tokyo, Japan, 25-29 July 2022)</em>, published by DH2022 Local Organizing Committee, 160‑65. Tokyo, 2022. https://dh2022.dhii.asia/dh2022bookofabsts.pdf#page=162 and https://dh2022.dhii.asia/abstracts/files/CAMPS_Jean_Baptiste_Data_Diversity_in_handwritten_text_recog.html. 本数据集包含以下目录: <strong>/texts/</strong>:存储主TEI文件,包含所有文本、语义及图形信息。 <strong>/img/</strong>:包含49幅底图,转录文本基于此制作,并通过TEI文件中的<facsimile/>元素实现文本与图像的行级对齐。 <strong>/alto/</strong>:为每幅图像生成对应ALTO格式文件,根据上述会议论文实验中的用户需求分为多个变体。文本为纯平格式,按行划分,支持扩写或保留缩写形式,同时可选择标准化处理或保留原始编码。鉴于Kraken/eScriptorium当前的实现逻辑,<strong>/alto/without-norm-coord-commas/</strong>目录下的坐标格式为`x1,y1 x2,y2...`,其余目录下的坐标格式则为`x1 y1 x2 y2...`。 <strong>/img-masks/</strong>:包含小型图像,展示TEI文件<facsimile/>元素中记录的所有行坐标,其中包含未被选入转录文本的行。



