Thai local dialect benchmark
收藏资源简介:
本研究引入了一个用于评估大型语言模型在处理泰国本地方言时的表现的新基准。该基准涵盖了北部的Lanna方言、东北部的Isan方言和南部的Dambro方言。数据集包含了针对总结、问答、翻译、对话和食物相关任务的样本,所有的输入、上下文、提示和标签都是由本地方言的母语者翻译的。该数据集旨在评估LLM对泰国本地方言的理解和生成能力。
This study introduces a novel benchmark for evaluating the performance of large language models (LLMs) when handling Thai regional dialects. The benchmark covers three primary dialects: Northern Thai (Lanna), Northeastern Thai (Isan), and Southern Thai (Dambro). The dataset includes samples for tasks such as summarization, question answering, translation, dialogue, and food-related tasks. All inputs, contexts, prompts, and labels were translated by native speakers of the respective local dialects. This benchmark aims to assess the understanding and generation capabilities of LLMs towards Thai regional dialects.




