MDMIC - Indic Cross-Domain, Multi-Intent NLU Dataset
收藏资源简介:
We construct an Indic benchmark corpus MDMIC through data augmentation, based on the benchmark multilingual dataset MASSIVE, for NLU tasks, i.e, ID, DC, SF in complex utterances in low-resource languages. These utterances span multiple domains and intents, comprising complex, multi-sentence structures.Each class in this dataset is a combination of multiple intents and domains. 6 dataset files are uplaoded as per the given language <Hindi, Bengali, Tamil, Telugu, Kannada, Malayalam>. It has Complex Sentence (contains the user-utterance) , Intents (refers to the multiple intents in the user-utterance) , Scenarios (refers to the domain information - contains cross domain data) , annot_uts (contains slot\/entity level information in the specified language), num_scenarios (contain the number of domains), annot_utts_eng (contains the slot\/entity level information in English)



