xSID 0.5, de-muc
收藏资源简介:
xSID 0.5数据集由慕尼黑大学信息与语言处理中心发布,包含44000条句子,涵盖16种意图和33种槽位类型,数据来源于SNIPS和Facebook的重新标注句子。de-muc数据集是作者新发布的慕尼黑巴伐利亚方言数据集,用于评估方言变体对模型性能的影响。数据集通过翻译和标注生成,反映了方言的拼写和语法特征。这些数据集主要用于自然语言理解中的槽位和意图检测任务,旨在解决方言数据稀缺和模型在方言数据上表现不佳的问题。
The xSID 0.5 dataset, released by the Center for Information and Language Processing at LMU Munich, contains 44,000 sentences covering 16 intent categories and 33 slot types, with its data sourced from re-annotated sentences from SNIPS and Facebook. The de-muc dataset is a newly released Bavarian dialect dataset from Munich developed by the authors, intended to evaluate the impact of dialectal variations on model performance. This dataset is constructed through translation and manual annotation, capturing the spelling and grammatical features of the dialect. Both datasets are primarily utilized for slot and intent detection tasks in Natural Language Understanding (NLU), aiming to address the scarcity of dialectal data and the subpar performance of models on dialectal datasets.




