The Knesset Meetings Corpus 2004-2005
收藏资源简介:
The Knesset Meetings Corpus 2004-2005 is made up of two components: Raw texts - 282 files made up of 867,725 lines together. These can be downloaded in two formats: As <code>doc</code> files, encoded using <code>windows-1255</code> encoding: <code>kneset16.zip</code> - Contains 164 text files made up of 543,228 lines together. [MILA host] [Github Mirror] <code>kneset17.zip</code> - Contains 118 text files made up of 324,497 lines together. [MILA host] [Github Mirror] As <code>txt</code> files, encoded using <code>utf8</code> encoding: <code>kneset.tar.gz</code> - An archive of all the raw text files, divided into two folders: [Github mirror] <code>16</code> - Contains 164 text files made up of 543,228 lines together. <code>17</code> - Contains 118 text files made up of 324,497 lines together. <code>knesset_txt_16.tar.gz</code>- Contains 164 text files made up of 543,228 lines together. [MILA host] [Github Mirror] <code>knesset_txt_17.zip</code> - Contains 118 text files made up of 324,497 lines together. [MILA host] [Github Mirror] Tokenized and morphologically tagged texts - Tagged versions exist only for the files in the <code>16</code> folder. The text are represented using MILA's XML schema for corpora. These can be downloaded in two ways: <code>knesset_tagged_16.tar.gz</code> - An archive of all tokenized and tagged files. [MILA host] [Archive.org mirror] By cloning this repository, as the unarchived version of these files can be found in this repository, under the <code>knesset_tagged</code> folder.
2004-2005年以色列议会会议语料库(Knesset Meetings Corpus 2004-2005)由两个部分构成: 一、原始文本:共计282个文件,总行数达867,725行。该原始文本提供两种下载格式: 1. `<code>doc</code>格式文件,采用`<code>windows-1255</code>`编码: - `<code>kneset16.zip</code>`:包含164个文本文件,总行数为543,228行。[MILA主机镜像] [GitHub镜像站] - `<code>kneset17.zip</code>`:包含118个文本文件,总行数为324,497行。[MILA主机镜像] [GitHub镜像站] 2. `<code>txt</code>格式文件,采用`<code>utf8</code>`编码: - `<code>kneset.tar.gz</code>`:包含全部原始文本文件的归档包,分为两个子文件夹:[GitHub镜像站] - `<code>16</code>`:包含164个文本文件,总行数为543,228行。 - `<code>17</code>`:包含118个文本文件,总行数为324,497行。 - `<code>knesset_txt_16.tar.gz</code>`:包含164个文本文件,总行数为543,228行。[MILA主机镜像] [GitHub镜像站] - `<code>knesset_txt_17.zip</code>`:包含118个文本文件,总行数为324,497行。[MILA主机镜像] [GitHub镜像站] 二、分词与形态标注文本:仅针对`<code>16</code>`文件夹内的文件提供标注版本。此类文本采用MILA面向语料库制定的XML模式(XML Schema)进行表征,提供两种下载方式: 1. `<code>knesset_tagged_16.tar.gz</code>`:包含全部分词标注文件的归档包。[MILA主机镜像] [Archive.org镜像站] 2. 通过克隆该代码仓库获取:该仓库的`<code>knesset_tagged</code>`文件夹下包含上述未归档的标注文件。



