Spider, BIRD
收藏资源简介:
本研究涉及的两个主要数据集为Spider和BIRD。Spider数据集包含10,181个自然语言查询,涵盖200个数据库和138个领域,分为四个难度级别。BIRD数据集则包含12,751个问题-SQL对,涉及95个大型数据库,分为三个难度级别。这两个数据集均用于评估和调优大型语言模型在文本到SQL任务中的性能。创建过程中,数据集通过标准化流程处理,确保数据质量和适用性。应用领域主要集中在自然语言处理和数据库交互,旨在提高非专业用户与数据库系统之间的交互效率,解决复杂的查询生成问题。
The two core datasets utilized in this study are Spider and BIRD. The Spider dataset comprises 10,181 natural language queries, spanning 200 databases across 138 domains, and is divided into four difficulty tiers. The BIRD dataset contains 12,751 question-SQL pairs, involving 95 large-scale databases, and is categorized into three difficulty levels. Both datasets are employed to evaluate and fine-tune the performance of Large Language Models (LLMs) on the text-to-SQL task. During their development, the datasets were processed through standardized procedures to guarantee data quality and applicability. Their application domains primarily center on natural language processing (NLP) and database interaction, aiming to improve the interaction efficiency between non-professional users and database systems, and address complex query generation problems.




