遇见数据集

Token-Level Post-Editing Dataset (EN–MT–Human): English-Ukrainian Translation Edit Log (Education-Legal)

收藏
Zenodo2026-02-23 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains a complete token-level log of post-editing operations (N = 6208 edits) extracted from a triple-aligned corpus consisting of: English source text DeepL machine translation (Ukrainian) Human-edited Ukrainian translation Sentence alignment was performed using a length-based Gale–Church alignment variant. Token-level differences between the machine translation (MT) and the human-edited version were extracted using sequence-based comparison (replace / insert / delete operations). The dataset records all non-equal token-level modifications introduced during human post-editing of machine translation. The dataset includes: Unique edit identifiers Alignment block identifiers Page references to the English source Token-level edit operations Rule-based edit categorization This version does not contain full textual contexts of the source, MT, or human-edited translation. It provides the complete analytical layer of post-editing operations without redistributing copyrighted textual material. The dataset is intended for research in post-editing studies, translation process research, machine translation evaluation, and corpus-based translation analysis.

提供机构:
Zenodo
创建时间:
2026-02-23
二维码
社区交流群
二维码
科研交流群
商业服务