ytu-ce-cosmos/absa-tr
收藏资源简介:
ABSA-TR是一个土耳其语方面级情感分析数据集,包含来自电子商务、补充剂和电影领域的16,031条真实用户评论句子和24,439个方面标注。数据集支持方面极性分类和方面术语提取任务,其中极性分类包括隐式方面,而提取任务仅使用显式跨度。数据划分为训练集(11,993个句子,17,423个方面,3,995个隐式方面)、验证集(688个句子,1,218个方面,179个隐式方面)和测试集(3,350个句子,5,798个方面,804个隐式方面)。每条数据包含pool_id、domain、text、aspects和flags列,每个方面有字面span、归一化aspect、极性(正面、负面或中性)和标注source。训练集由DeepSeek-V4-Pro标注,验证集和测试集基于GLM-5.1、MiniMax-M3和Kimi-K2.6的共识标注,数据源自Altinok (2023)介绍的土耳其评论资源。
A Turkish aspect-based sentiment dataset with 16,031 real user-review sentences and 24,439 aspect annotations from e-commerce, supplements, and movie domains. The dataset supports aspect polarity classification and aspect-term extraction, with polarity classification including implicit aspects and extraction using only explicit spans. It is split into train (11,993 sentences, 17,423 aspects, 3,995 implicit aspects), validation (688 sentences, 1,218 aspects, 179 implicit aspects), and test (3,350 sentences, 5,798 aspects, 804 implicit aspects). Each row contains pool_id, domain, text, aspects, and flags, with each aspect having a verbatim span, a normalized aspect, a polarity of positive, negative, or neutral, and an annotation source. The training split was annotated by DeepSeek-V4-Pro, while validation and test use a separate pool labeled by consensus among GLM-5.1, MiniMax-M3, and Kimi-K2.6, with original sentences from Turkish review resources introduced by Altinok (2023).




