Token-Level Post-Editing Dataset (EN–MT–Human): English-Ukrainian Translation Edit Log (Education-Legal)
收藏资源简介:
This dataset contains a complete token-level log of post-editing operations (N = 6208 edits) extracted from a triple-aligned corpus consisting of: English source text DeepL machine translation (Ukrainian) Human-edited Ukrainian translation Sentence alignment was performed using a length-based Gale–Church alignment variant. Token-level differences between the machine translation (MT) and the human-edited version were extracted using sequence-based comparison (replace / insert / delete operations). The dataset records all non-equal token-level modifications introduced during human post-editing of machine translation. The dataset includes: Unique edit identifiers Alignment block identifiers Page references to the English source Token-level edit operations Rule-based edit categorization This version does not contain full textual contexts of the source, MT, or human-edited translation. It provides the complete analytical layer of post-editing operations without redistributing copyrighted textual material. The dataset is intended for research in post-editing studies, translation process research, machine translation evaluation, and corpus-based translation analysis.



