usc-isi/hansards
收藏资源简介:
该版本包含来自加拿大第36届议会官方记录(Hansards)的130万对对齐文本块(句子或更小的片段)。完整的Hansards记录包括众议院和参议院的辩论,尽可能对齐后,语料库被分为五组句子对:训练集(占句子对的80%)、两组测试集(各占5%)和两组最终评估集(各占5%)。当前版本包含训练集和测试集,评估集保留用于未来的机器翻译评估目的,目前不可用。需要注意的是,该版本仅包含句子对,可能存在由于多对一、多对多或一对多对齐被过滤掉而导致的句子顺序不一致,因此可能不适用于与话语相关的研究。此外,句子分割和对齐并不完美,特别是长度差异较大的句子对,可能需要在统计训练前进行过滤。
This version contains 1.3 million aligned text chunks (sentences or shorter segments) from the official records of the 36th Canadian Parliament (Hansards). The complete Hansards corpus covers debates from both the House of Commons and the Senate. After aligning the records as thoroughly as possible, the corpus was split into five sets of sentence pairs: a training set (80% of all pairs), two test sets (5% each), and two final evaluation sets (5% each). This release only includes the training and test sets; the evaluation sets are reserved for future machine translation evaluation and are not currently available. Notably, this version only contains sentence pairs, and sentence order inconsistencies may arise due to the filtering of one-to-many, many-to-one, or many-to-many alignments. As such, it may not be suitable for discourse-related research. Additionally, sentence segmentation and alignment are not perfect, particularly for sentence pairs with large length discrepancies, and filtering may be required prior to statistical training.
数据集卡片:hansards
数据集概述
数据集摘要
该数据集包含130万对来自第36届加拿大议会官方记录(Hansards)的对齐文本块(句子或更小的片段)。数据集分为训练和测试集,其中训练集占80%,测试集占10%。评估集目前不可用,保留用于未来的机器翻译评估。
数据集结构
数据实例
house
- 下载的数据文件大小: 67.58 MB
- 生成的数据集大小: 214.37 MB
- 总磁盘使用量: 281.95 MB
训练集示例: json { "en": "Mr. Walt Lastewka (Parliamentary Secretary to Minister of Industry, Lib.):", "fr": "M. Walt Lastewka (secrétaire parlementaire du ministre de lIndustrie, Lib.):" }
senate
- 下载的数据文件大小: 15.25 MB
- 生成的数据集大小: 46.03 MB
- 总磁盘使用量: 61.28 MB
训练集示例: json { "en": "Mr. Walt Lastewka (Parliamentary Secretary to Minister of Industry, Lib.):", "fr": "M. Walt Lastewka (secrétaire parlementaire du ministre de lIndustrie, Lib.):" }
数据字段
house
fr: 字符串类型特征。en: 字符串类型特征。
senate
fr: 字符串类型特征。en: 字符串类型特征。
数据分割
| 名称 | 训练集 | 测试集 |
|---|---|---|
| house | 947969 | 122290 |
| senate | 182135 | 25553 |
数据集创建
数据集配置
senate
- 特征:
fr: 字符串类型en: 字符串类型
- 分割:
test: 5711686字节,25553个样本train: 40324278字节,182135个样本
- 下载大小: 15247360字节
- 数据集大小: 46035964字节
house
- 特征:
fr: 字符串类型en: 字符串类型
- 分割:
test: 22906629字节,122290个样本train: 191459584字节,947969个样本
- 下载大小: 67584000字节
- 数据集大小: 214366213字节




