Dialect2SQL
收藏资源简介:
Dialect2SQL是由穆罕默德六世理工大学和法赫德国王石油矿产大学的研究团队创建的首个针对摩洛哥方言的大规模跨领域文本到SQL数据集。该数据集包含9428个自然语言问题与SQL查询对,覆盖69个不同领域的数据库,涉及食品、书籍、教育、交通、犯罪等多个领域。数据集的创建过程包括使用GPT-4进行自动翻译,并由母语为摩洛哥方言的计算机科学学生进行人工校对,以确保翻译质量。该数据集旨在解决低资源语言在文本到SQL任务中的挑战,特别是摩洛哥方言的复杂性,如多样的词汇来源、借词和独特的表达方式。Dialect2SQL的应用领域包括自然语言处理、数据库查询生成以及低资源语言的研究与开发。
Dialect2SQL is the first large-scale cross-domain text-to-SQL dataset tailored for Moroccan dialect, developed by research teams from Université Mohammed VI Polytechnique and King Fahd University of Petroleum and Minerals. It consists of 9,428 paired samples of natural language questions and corresponding SQL queries, covering databases across 69 distinct domains such as food, books, education, transportation, crime and more. The dataset was constructed through automatic translation powered by GPT-4, followed by manual verification conducted by computer science students who are native Moroccan dialect speakers to ensure translation quality. This dataset aims to tackle the challenges faced by low-resource languages in text-to-SQL tasks, particularly the inherent complexity of Moroccan dialect including diverse lexical origins, loanwords and distinctive expressive patterns. The application fields of Dialect2SQL include natural language processing, database query generation, as well as research and development of low-resource languages.




