CARBD-Ko
收藏资源简介:
CARBD-Ko是一个专为韩语设计的上下文标注评论基准数据集,用于方面级情感分类。该数据集由首尔国立大学创建,包含10586个训练实例,旨在通过特定方面、方面极性、方面无关极性和方面强度等标注,解决预训练语言模型中的上下文化和幻觉问题。数据集的创建过程包括从多个领域收集公开评论、提取方面和意见、手动标注极性和强度,并通过同行评审确保标注的客观性和准确性。CARBD-Ko的应用领域主要集中在方面级情感分类,特别是解决模型在不同上下文中准确预测情感极性的挑战。
CARBD-Ko is a context-annotated review benchmark dataset tailored specifically for Korean, designed for aspect-based sentiment classification. Developed by Seoul National University, this dataset comprises 10,586 training instances. Its core objective is to mitigate the contextualization and hallucination issues prevalent in pretrained language models, with annotations covering specific aspects, aspect polarities, aspect-agnostic polarities, and aspect intensities. The dataset construction workflow includes collecting public reviews from diverse domains, extracting aspects and associated opinions, manually annotating polarities and intensity levels, and conducting peer reviews to ensure the objectivity and accuracy of the annotation results. The primary application scenarios of CARBD-Ko center on aspect-based sentiment classification, particularly addressing the challenge of enabling models to accurately predict sentiment polarities across different contextual settings.




