Helsinki-NLP/kde4
收藏资源简介:
--- annotations_creators: - found language_creators: - found language: - af - ar - as - ast - be - bg - bn - br - ca - crh - cs - csb - cy - da - de - el - en - eo - es - et - eu - fa - fi - fr - fy - ga - gl - gu - ha - he - hi - hne - hr - hsb - hu - hy - id - is - it - ja - ka - kk - km - kn - ko - ku - lb - lt - lv - mai - mk - ml - mr - ms - mt - nb - nds - ne - nl - nn - nso - oc - or - pa - pl - ps - pt - ro - ru - rw - se - si - sk - sl - sr - sv - ta - te - tg - th - tr - uk - uz - vi - wa - xh - zh language_bcp47: - bn-IN - en-GB - pt-BR - zh-CN - zh-HK - zh-TW license: - unknown multilinguality: - multilingual size_categories: - 100K<n<1M source_datasets: - original task_categories: - translation task_ids: [] paperswithcode_id: null pretty_name: KDE4 dataset_info: - config_name: fi-nl features: - name: id dtype: string - name: translation dtype: translation: languages: - fi - nl splits: - name: train num_bytes: 8845933 num_examples: 101593 download_size: 2471355 dataset_size: 8845933 - config_name: it-ro features: - name: id dtype: string - name: translation dtype: translation: languages: - it - ro splits: - name: train num_bytes: 8827049 num_examples: 109003 download_size: 2389051 dataset_size: 8827049 - config_name: nl-sv features: - name: id dtype: string - name: translation dtype: translation: languages: - nl - sv splits: - name: train num_bytes: 22294586 num_examples: 188454 download_size: 6203460 dataset_size: 22294586 - config_name: en-it features: - name: id dtype: string - name: translation dtype: translation: languages: - en - it splits: - name: train num_bytes: 27132585 num_examples: 220566 download_size: 7622662 dataset_size: 27132585 - config_name: en-fr features: - name: id dtype: string - name: translation dtype: translation: languages: - en - fr splits: - name: train num_bytes: 25650409 num_examples: 210173 download_size: 7049364 dataset_size: 25650409 --- # Dataset Card for KDE4 ## Table of Contents - [Dataset Description](#dataset-description) - [Dataset Summary](#dataset-summary) - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards) - [Languages](#languages) - [Dataset Structure](#dataset-structure) - [Data Instances](#data-instances) - [Data Fields](#data-fields) - [Data Splits](#data-splits) - [Dataset Creation](#dataset-creation) - [Curation Rationale](#curation-rationale) - [Source Data](#source-data) - [Annotations](#annotations) - [Personal and Sensitive Information](#personal-and-sensitive-information) - [Considerations for Using the Data](#considerations-for-using-the-data) - [Social Impact of Dataset](#social-impact-of-dataset) - [Discussion of Biases](#discussion-of-biases) - [Other Known Limitations](#other-known-limitations) - [Additional Information](#additional-information) - [Dataset Curators](#dataset-curators) - [Licensing Information](#licensing-information) - [Citation Information](#citation-information) - [Contributions](#contributions) ## Dataset Description - **Homepage:** http://opus.nlpl.eu/KDE4.php - **Repository:** None - **Paper:** http://www.lrec-conf.org/proceedings/lrec2012/pdf/463_Paper.pdf - **Leaderboard:** [More Information Needed] - **Point of Contact:** [More Information Needed] ### Dataset Summary To load a language pair which isn't part of the config, all you need to do is specify the language code as pairs. You can find the valid pairs in Homepage section of Dataset Description: http://opus.nlpl.eu/KDE4.php E.g. `dataset = load_dataset("kde4", lang1="en", lang2="nl")` ### Supported Tasks and Leaderboards [More Information Needed] ### Languages [More Information Needed] ## Dataset Structure ### Data Instances [More Information Needed] ### Data Fields [More Information Needed] ### Data Splits [More Information Needed] ## Dataset Creation ### Curation Rationale [More Information Needed] ### Source Data [More Information Needed] #### Initial Data Collection and Normalization [More Information Needed] #### Who are the source language producers? [More Information Needed] ### Annotations [More Information Needed] #### Annotation process [More Information Needed] #### Who are the annotators? [More Information Needed] ### Personal and Sensitive Information [More Information Needed] ## Considerations for Using the Data ### Social Impact of Dataset [More Information Needed] ### Discussion of Biases [More Information Needed] ### Other Known Limitations [More Information Needed] ## Additional Information ### Dataset Curators [More Information Needed] ### Licensing Information [More Information Needed] ### Citation Information [More Information Needed] ### Contributions Thanks to [@abhishekkrthakur](https://github.com/abhishekkrthakur) for adding this dataset.
annotations_creators: - 标注创建者 language_creators: - 语言创建者 language: - 南非语(af) - 阿拉伯语(ar) - 阿萨姆语(as) - 阿斯图里亚斯语(ast) - 白俄罗斯语(be) - 保加利亚语(bg) - 孟加拉语(bn) - 布列塔尼语(br) - 加泰罗尼亚语(ca) - 克里米亚鞑靼语(crh) - 捷克语(cs) - 卡舒比语(csb) - 威尔士语(cy) - 丹麦语(da) - 德语(de) - 希腊语(el) - 英语(en) - 世界语(eo) - 西班牙语(es) - 爱沙尼亚语(et) - 巴斯克语(eu) - 波斯语(fa) - 芬兰语(fi) - 法语(fr) - 弗里西语(fy) - 爱尔兰语(ga) - 加利西亚语(gl) - 古吉拉特语(gu) - 豪萨语(ha) - 希伯来语(he) - 印地语(hi) - 切蒂斯格尔语(hne) - 克罗地亚语(hr) - 上索布语(hsb) - 匈牙利语(hu) - 亚美尼亚语(hy) - 印尼语(id) - 冰岛语(is) - 意大利语(it) - 日语(ja) - 格鲁吉亚语(ka) - 哈萨克语(kk) - 高棉语(km) - 卡纳达语(kn) - 韩语(ko) - 库尔德语(ku) - 卢森堡语(lb) - 立陶宛语(lt) - 拉脱维亚语(lv) - 迈蒂利语(mai) - 马其顿语(mk) - 马拉雅拉姆语(ml) - 马拉地语(mr) - 马来语(ms) - 马耳他语(mt) - 挪威博克马尔语(nb) - 低地德语(nds) - 尼泊尔语(ne) - 荷兰语(nl) - 挪威尼诺斯克语(nn) - 北索托语(nso) - 奥克语(oc) - 奥里亚语(or) - 旁遮普语(pa) - 波兰语(pl) - 普什图语(ps) - 葡萄牙语(pt) - 罗马尼亚语(ro) - 俄语(ru) - 卢旺达语(rw) - 北萨米语(se) - 僧伽罗语(si) - 斯洛伐克语(sk) - 斯洛文尼亚语(sl) - 塞尔维亚语(sr) - 瑞典语(sv) - 泰米尔语(ta) - 泰卢固语(te) - 塔吉克语(tg) - 泰语(th) - 土耳其语(tr) - 乌克兰语(uk) - 乌兹别克语(uz) - 越南语(vi) - 瓦隆语(wa) - 科萨语(xh) - 中文(zh) language_bcp47: - 孟加拉语(印度)(bn-IN) - 英语(英国)(en-GB) - 葡萄牙语(巴西)(pt-BR) - 简体中文(zh-CN) - 香港繁体中文(zh-HK) - 台湾繁体中文(zh-TW) license: - 未知 multilinguality: - 多语言 size_categories: - 10万<样本数<100万 source_datasets: - 原创数据集 task_categories: - 机器翻译 task_ids: [] paperswithcode_id: null pretty_name: KDE4 dataset_info: - config_name: 芬兰语-荷兰语(fi-nl) features: - name: id dtype: 字符串 - name: translation dtype: translation: languages: - fi - nl splits: - name: train num_bytes: 8845933 num_examples: 101593 download_size: 2471355 dataset_size: 8845933 - config_name: 意大利语-罗马尼亚语(it-ro) features: - name: id dtype: 字符串 - name: translation dtype: translation: languages: - it - ro splits: - name: train num_bytes: 8827049 num_examples: 109003 download_size: 2389051 dataset_size: 8827049 - config_name: 荷兰语-瑞典语(nl-sv) features: - name: id dtype: 字符串 - name: translation dtype: translation: languages: - nl - sv splits: - name: train num_bytes: 22294586 num_examples: 188454 download_size: 6203460 dataset_size: 22294586 - config_name: 英语-意大利语(en-it) features: - name: id dtype: 字符串 - name: translation dtype: translation: languages: - en - it splits: - name: train num_bytes: 27132585 num_examples: 220566 download_size: 7622662 dataset_size: 27132585 - config_name: 英语-法语(en-fr) features: - name: id dtype: 字符串 - name: translation dtype: translation: languages: - en - fr splits: - name: train num_bytes: 25650409 num_examples: 210173 download_size: 7049364 dataset_size: 25650409 # KDE4数据集卡片 ## 目录 - [数据集描述](#dataset-description) - [数据集摘要](#dataset-summary) - [支持任务与排行榜](#supported-tasks-and-leaderboards) - [语言](#languages) - [数据集结构](#dataset-structure) - [数据实例](#data-instances) - [数据字段](#data-fields) - [数据划分](#data-splits) - [数据集创建](#dataset-creation) - [筛选依据](#curation-rationale) - [源数据](#source-data) - [标注](#annotations) - [个人与敏感信息](#personal-and-sensitive-information) - [数据集使用注意事项](#considerations-for-using-the-data) - [数据集的社会影响](#social-impact-of-dataset) - [偏差讨论](#discussion-of-biases) - [其他已知限制](#other-known-limitations) - [附加信息](#additional-information) - [数据集整理者](#dataset-curators) - [许可信息](#licensing-information) - [引用信息](#citation-information) - [贡献](#contributions) ## 数据集描述 - **主页**:http://opus.nlpl.eu/KDE4.php - **仓库**:无 - **论文**:http://www.lrec-conf.org/proceedings/lrec2012/pdf/463_Paper.pdf - **排行榜**:[需补充更多信息] - **联系人**:[需补充更多信息] ### 数据集摘要 若需加载配置中未涵盖的语言对,仅需指定双语言代码即可。有效语言对组合可在数据集描述的主页链接中查询:http://opus.nlpl.eu/KDE4.php。示例代码如下: `dataset = load_dataset("kde4", lang1="en", lang2="nl")` ### 支持任务与排行榜 [需补充更多信息] ### 语言 [需补充更多信息] ## 数据集结构 ### 数据实例 [需补充更多信息] ### 数据字段 [需补充更多信息] ### 数据划分 [需补充更多信息] ## 数据集创建 ### 筛选依据 [需补充更多信息] ### 源数据 [需补充更多信息] #### 初始数据收集与标准化 [需补充更多信息] #### 源语言创作者是谁? [需补充更多信息] ### 标注 [需补充更多信息] #### 标注流程 [需补充更多信息] #### 标注者是谁? [需补充更多信息] ### 个人与敏感信息 [需补充更多信息] ## 数据集使用注意事项 ### 数据集的社会影响 [需补充更多信息] ### 偏差讨论 [需补充更多信息] ### 其他已知限制 [需补充更多信息] ## 附加信息 ### 数据集整理者 [需补充更多信息] ### 许可信息 [需补充更多信息] ### 引用信息 [需补充更多信息] ### 贡献 感谢 [@abhishekkrthakur](https://github.com/abhishekkrthakur) 为本数据集添加支持。
数据集概述
- 名称: KDE4
- 语言: 多语言,包括但不限于 af, ar, as, ast, be, bg, bn, br, ca, crh, cs, csb, cy, da, de, el, en, eo, es, et, eu, fa, fi, fr, fy, ga, gl, gu, ha, he, hi, hne, hr, hsb, hu, hy, id, is, it, ja, ka, kk, km, kn, ko, ku, lb, lt, lv, mai, mk, ml, mr, ms, mt, nb, nds, ne, nl, nn, nso, oc, or, pa, pl, ps, pt, ro, ru, rw, se, si, sk, sl, sr, sv, ta, te, tg, th, tr, uk, uz, vi, wa, xh, zh
- BCP47语言标签: bn-IN, en-GB, pt-BR, zh-CN, zh-HK, zh-TW
- 许可证: 未知
- 多语言性: 多语言
- 大小类别: 100K<n<1M
- 源数据集: 原始数据
- 任务类别: 翻译
数据集结构
- 配置名称: fi-nl
- 特征:
- id: 字符串类型
- translation: 翻译类型,包含语言 fi, nl
- 分割:
- train: 101593个示例,8845933字节,下载大小2471355字节
- 特征:
- 配置名称: it-ro
- 特征:
- id: 字符串类型
- translation: 翻译类型,包含语言 it, ro
- 分割:
- train: 109003个示例,8827049字节,下载大小2389051字节
- 特征:
- 配置名称: nl-sv
- 特征:
- id: 字符串类型
- translation: 翻译类型,包含语言 nl, sv
- 分割:
- train: 188454个示例,22294586字节,下载大小6203460字节
- 特征:
- 配置名称: en-it
- 特征:
- id: 字符串类型
- translation: 翻译类型,包含语言 en, it
- 分割:
- train: 220566个示例,27132585字节,下载大小7622662字节
- 特征:
- 配置名称: en-fr
- 特征:
- id: 字符串类型
- translation: 翻译类型,包含语言 en, fr
- 分割:
- train: 210173个示例,25650409字节,下载大小7049364字节
- 特征:




