Handbooks prepared by the Historical Section of the Foreign Office
收藏资源简介:
The PeaceBooks corpus is a set of volumes with historical, ethnographic, cartographic, and economic information on almost every part of the world prepared by the Historical Section of the Foreign Office for use by the British delegates to the Paris Peace Conference. The PeaceBooks corpus is composed of texts that were published in 25 volumes in 1920. The copies that make up this corpus were digitized by Goole and were provided by the University of Iowa (16), University of Michigan (8), and one volume by the University of Wisconsin (volume 6 covered France, Italy, and Spain). The orginal files were downloaded from HathiTrust. Typographical errors and OCR mistakes were corrected. Different peritext elements, such as volume titles, chapter and section headings, footnotes, page numbers, and titles in text of all volumes, were annotated. The structural elements were tagged by one or multiple hash signs '#', reflecting the hierarchicals structure of the volume, footnotes start with tilda sign '~'. The work is in public domain. It was idigitized and OCR-ed by Google. The file peace_books_ids.tsv contains the identifiers for the source files in HathiTrust, as well as metadata for the individual files (explanation of the columns are in the related GitHub repository).
PeaceBooks语料库(PeaceBooks corpus)是由英国外交部历史司编制的一套卷宗,涵盖全球几乎所有地区的历史、民族志、地图学与经济信息,专供出席巴黎和会的英国代表团使用。 该语料库由1920年出版的25卷文献构成。构成该语料库的数字化副本由谷歌(Google)完成,并由爱荷华大学(University of Iowa,16卷)、密歇根大学(University of Michigan,8卷)以及威斯康星大学(University of Wisconsin,第6卷内容涵盖法国、意大利与西班牙)提供1卷。 原始文件均从HathiTrust平台下载获取。文本排版错误与OCR识别误差均已修正。所有卷宗的各类副文本(peritext)元素,如卷名、章节标题、脚注(footnote)、页码以及正文中的标题,均已完成标注。结构元素通过单个或多个井号(#)进行标记,以体现卷宗的层级结构;脚注则以波浪号(~)作为起始标识。 该作品已进入公有领域,其数字化与OCR处理均由谷歌(Google)完成。 peace_books_ids.tsv文件包含HathiTrust平台上源文件的标识符,以及各单个文件的元数据(metadata),各列的详细说明详见相关GitHub仓库。



