韩国句子结尾数据集(KoSEnd)
收藏资源简介:
韩国句子结尾数据集(KoSEnd)是一个包含3000个句子的数据集,每个句子都标注了15种句子结尾形式。这些句子从不同的来源收集而来,涵盖了各种语境。该数据集旨在评估大型语言模型(LLMs)对韩语句子的理解能力,特别是对复杂句子结尾的理解。数据集的构建过程包括语料库收集、句子结尾扩展和两阶段标注。研究结果表明,LLMs在处理韩语句子结尾时存在一定的挑战,但通过引入句子结尾可能缺失的概念,模型的性能得到了显著提升。
The Korean Sentence End Dataset (KoSEnd) is a dataset containing 3,000 sentences, each annotated with 15 types of sentence endings. These sentences are collected from diverse sources and cover a wide range of contexts. This dataset is designed to evaluate the ability of Large Language Models (LLMs) to understand Korean sentences, particularly complex sentence endings. The construction process of this dataset includes corpus collection, sentence ending expansion, and two-stage annotation. Research findings indicate that LLMs face certain challenges when processing Korean sentence endings; however, introducing the concept of potentially missing sentence endings leads to a significant improvement in model performance.




