SEMANTIC ERRORS IN MACHINE TRANSLATION SYSTEMS AND THEIR CAUSES
收藏资源简介:
Semantic errors in machine translation (MT) occur when the output is fluent but fails to preserve meaning—by mistranslating word senses, omitting or adding content, or generating “hallucinated” information not supported by the source. Such errors are especially problematic because they can look grammatically perfect while being semantically wrong, making them hard to detect during post-editing. Research on neural machine translation (NMT) highlights that adequacy problems such as omissions and additions can appear in otherwise fluent output, masking meaning loss. Studies on hallucinations show that NMT can produce translations “untethered” from the input, sometimes triggered by rare tokens or distribution shifts. This article explains the main types of semantic MT errors, links them to underlying causes (lexical ambiguity, domain shift, data noise, decoding behavior, and model uncertainty), and illustrates them with short examples followed by clarifying commentary. It also summarizes evaluation practices (MQM-style categories and adequacy-focused learned metrics) and argues for targeted quality checks beyond surface fluency.



