遇见数据集

The Knesset Meetings Corpus 2004-2005

收藏
Zenodo2020-07-29 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

The Knesset Meetings Corpus 2004-2005 is made up of two components: Raw texts - 282 files made up of 867,725 lines together. These can be downloaded in two formats: As <code>doc</code> files, encoded using <code>windows-1255</code> encoding: <code>kneset16.zip</code> - Contains 164 text files made up of 543,228 lines together. [MILA host] [Github Mirror] <code>kneset17.zip</code> - Contains 118 text files made up of 324,497 lines together. [MILA host] [Github Mirror] As <code>txt</code> files, encoded using <code>utf8</code> encoding: <code>kneset.tar.gz</code> - An archive of all the raw text files, divided into two folders: [Github mirror] <code>16</code> - Contains 164 text files made up of 543,228 lines together. <code>17</code> - Contains 118 text files made up of 324,497 lines together. <code>knesset_txt_16.tar.gz</code>- Contains 164 text files made up of 543,228 lines together. [MILA host] [Github Mirror] <code>knesset_txt_17.zip</code> - Contains 118 text files made up of 324,497 lines together. [MILA host] [Github Mirror] Tokenized and morphologically tagged texts - Tagged versions exist only for the files in the <code>16</code> folder. The text are represented using MILA's XML schema for corpora. These can be downloaded in two ways: <code>knesset_tagged_16.tar.gz</code> - An archive of all tokenized and tagged files. [MILA host] [Archive.org mirror] By cloning this repository, as the unarchived version of these files can be found in this repository, under the <code>knesset_tagged</code> folder.

提供机构:
Zenodo
创建时间:
2019-05-10
二维码
社区交流群
二维码
科研交流群
商业服务