DeepPavlov/Massive_es
收藏资源简介:
该数据集是一个多语言任务导向对话数据集,包含多个场景(如社交、交通、日历、播放、新闻等)和意图(如日期查询、物联网控制、交通票务、外卖查询等)。数据特征包括id、locale、partition、scenario、intent、utt(原始话语)、annot_utt(注释话语)、worker_id、slot_method(槽位和方法结构)和judgments(评分结构,包括意图得分、槽位得分等)。数据集分为训练集(11514个示例)、验证集(2033个示例)和测试集(2974个示例),用于自然语言处理任务,如意图分类和槽位填充。
This dataset is a multilingual task-oriented dialogue dataset encompassing multiple scenarios (e.g., social, transport, calendar, play, news) and intents (e.g., datetime_query, iot_hue_lightchange, transport_ticket, takeaway_query). Features include id, locale, partition, scenario, intent, utt (original utterance), annot_utt (annotated utterance), worker_id, slot_method (a structured field with slot and method lists), and judgments (a structured field with worker_id, intent_score, slots_score, etc.). It is split into train (11,514 examples), validation (2,033 examples), and test (2,974 examples) sets, designed for natural language processing tasks such as intent classification and slot filling.




