Code-Mixed Goal Oriented Conversation Systems Dataset
收藏资源简介:
本数据集名为Code-Mixed Goal Oriented Conversation Systems Dataset,由印度理工学院马德拉斯分校创建,旨在支持多语言和代码混合对话系统的发展。数据集包含49167条对话,涵盖了印度四种主要语言(印地语、孟加拉语、古吉拉特语和泰米尔语)与英语的代码混合对话。创建过程中,研究人员采用了内部和众包工作者的混合方式,确保对话的自然性和准确性。该数据集主要用于研究和开发能够理解和生成代码混合语言的对话系统,以满足多语言地区用户的需求。
This dataset, named Code-Mixed Goal Oriented Conversation Systems Dataset, was created by the Indian Institute of Technology Madras, aiming to support the development of multilingual and code-mixed conversational systems. It contains 49,167 dialogues, covering code-mixed conversations between English and four major Indian languages: Hindi, Bengali, Gujarati, and Tamil. During the dataset creation process, researchers adopted a hybrid approach combining internal work and crowdsourced workers to ensure the naturalness and accuracy of the dialogues. This dataset is mainly used for researching and developing conversational systems that can understand and generate code-mixed languages, so as to meet the needs of users in multilingual regions.




