Using ChatGPT and Other AI Engines to Vocalize Medieval Hebrew
收藏资源简介:
Hebrew is usually written without vowel points, making it challenging for some readers to decipher. This is especially true of medieval Hebrew, which can have nonstandard grammar and orthography. This paper tested four artificial intelligence (AI) tools by asking them to add vowel points to an unpublished medieval Hebrew translation of the Lord’s Prayer. The vocalization tools tested were OpenAI’s ChatGPT-3.5 and ChatGPT-4, Pellaworks’ DoItInHebrew, and Dicta’s Nakdan. ChatGPT-3.5 freely changed the text, even rewriting some phrases and adding an entire sentence. ChatGPT-3.5 also provided erroneous vowels in its rewritten Hebrew text. ChatGPT-4 did a moderately good job with only a few errors, but also modified the orthography. One of ChatGPT-4’s errors was not trivial, resulting in the invention of a word. When challenged, ChatGPT-4 corrected this confabulation by inventing another word, which it claimed was a “rare form” for which it provided a fictitious derivation. When challenged on this second made-up word, ChatGPT-4 replaced the word from the input text with a word based on an entirely different root. DoItInHebrew inserted vowels that produced a gibberish text. In contrast, Dicta’s Nakdan provided near perfect vocalization, with only one genuine error, but like ChatGPT-4 it modified the orthography. ChatGPT-3.5, ChatGPT-4, and DoItInHebrew exhibited serious “hallucinations,” of both the “factual” and the “untruthful” varieties, typical of other AIs, making them counterproductive for vocalizing historic Hebrew texts. Nakdan can be a powerful tool but still requires someone with expertise in Hebrew grammar to verify and correct the vocalization. Nakdan’s interface simplified correcting the vocalization, although it required its user to have advanced knowledge of Hebrew.
希伯来语通常省略元音符号,这使得部分读者难以识读,这一点在中世纪希伯来语中尤为突出——这类文本往往存在非标准语法与正字法。本研究通过要求四款人工智能(AI)工具为一份未公开的中世纪希伯来语版《主祷文》译文添加元音符号,对其进行了测试。本次测试的元音标注工具分别为OpenAI的ChatGPT-3.5与ChatGPT-4、Pellaworks的DoItInHebrew,以及Dicta的Nakdan。ChatGPT-3.5会随意修改文本,甚至重写部分短语并新增完整句子,其生成的希伯来语文本中还存在错误的元音标注。ChatGPT-4表现尚可,仅存在少量错误,但同样修改了文本的正字法;其中一处错误并非无关紧要,甚至造出新词。当被指出该问题时,ChatGPT-4为掩盖其虚构内容,又造出了另一个词,并声称这是一种“罕见形式”,还为其编造了词源。当再次被质疑这一杜撰词汇时,ChatGPT-4将输入文本中的原词替换为了基于完全不同词根的词汇。DoItInHebrew添加的元音符号会生成文理不通的乱码文本。与之形成对比的是,Dicta的Nakdan实现了近乎完美的元音标注,仅存在一处真实错误,但同样如ChatGPT-4一般修改了文本正字法。ChatGPT-3.5、ChatGPT-4与DoItInHebrew均出现了严重的“幻觉”问题,涵盖“事实性”与“虚假性”两类,这与其他人工智能系统的典型表现一致,因此它们用于标注古希伯来语文本的元音时反而会产生反效果。Nakdan虽可作为一款高效工具,但仍需要具备希伯来语语法专业知识的人员对其标注的元音进行校验与修正。尽管Nakdan的界面简化了元音标注的修正流程,但其使用仍要求用户掌握高阶希伯来语知识。



