遇见数据集

Spider-Realistic

收藏
arXiv2025-09-30 收录
官方服务:

资源简介:

该数据集名为Spider,包含了1034条英文表述及其对应的SQL查询语句,覆盖了20个不同的数据库架构。它旨在评估文本到SQL模型,允许研究人员提交模型预测的查询语句。此外,Spider数据集根据黄金SQL查询语句的复杂性,被分为四个难度级别:简单、中等、困难和超难。该数据集的规模为1034条表述,所涉及的任务是文本到SQL的转换。

The dataset named Spider contains 1034 pairs of English natural language utterances and their corresponding SQL query statements, spanning 20 distinct database schemas. It is developed to evaluate text-to-SQL models, and allows researchers to submit query statements predicted by their models. Additionally, the Spider dataset is divided into four difficulty tiers based on the complexity of its gold-standard SQL queries: simple, medium, hard, and extra hard. Comprising 1034 such utterance-SQL pairs, the core task of this dataset is text-to-SQL conversion.

提供机构:
Spider
搜集汇总
数据集介绍
Spider-Realistic 数据集图片
背景与挑战
背景概述
Spider-Realistic数据集是基于Spider数据集的开发集创建的变体,通过手动修改原始问题以移除显式列名提及,同时保持SQL查询不变,旨在更好地评估文本到SQL模型在自然语言与数据库模式对齐方面的能力。该数据集包含508个示例和19个数据库,是用于自然语言处理和数据库查询任务的重要评估资源。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务