bird_sql_dev_20251106
收藏资源简介:
BIRD-SQL Dev数据集是一个用于表格问答和问题回答任务的SQL数据集。该数据集经过了一个质量审查程序,由数据科学和AI领域的5位博士研究人员领导,并由具有10年以上经验的行业工程师和AI/数据科学领域的研究生团队支持。审查的目的是最小化模糊性并纠正错误,确保数据集的清晰性、一致性和可靠性。数据集中的问题由接受过BI训练的母语使用者编写。尽管努力减少了模糊性,但自然语言和NLP研究的固有特征仍然存在,这反映了在数据库上下文中解释人类问题的现实挑战。未来的排行榜将引入一个交互式的澄清设置,以帮助模型通过动态交互和澄清对话处理模糊性。
The BIRD-SQL Dev dataset is a SQL dataset intended for table-based question answering and general question answering tasks. This dataset underwent a quality review process led by five doctoral researchers in the fields of data science and artificial intelligence, and supported by a team of industry engineers with over 10 years of experience and graduate students specializing in AI and data science. The purpose of this review was to minimize ambiguity and correct errors, ensuring the clarity, consistency and reliability of the dataset. All questions in the dataset were written by native speakers trained in Business Intelligence (BI). Despite efforts to reduce ambiguity, the inherent characteristics of natural language and NLP research still exist, which reflects the real-world challenges of interpreting human queries within database contexts. Future leaderboards will introduce an interactive clarification framework to help models handle ambiguity through dynamic interactions and clarification dialogues.
BIRD-SQL Dev 数据集概述
基本信息
- 许可证: CC-BY-SA-4.0
- 任务类别: 表格问答、问答
- 语言: 英语
- 数据规模: 1K<n<10K
数据集描述
BIRD-SQL Dev 是一个用于文本到SQL转换任务的开发数据集,专注于数据库相关的自然语言问答。
数据集结构
每个数据条目包含以下字段:
| 字段名 | 数据类型 | 描述 |
|---|---|---|
question_id |
整数 | 每个实例的唯一标识符 |
db_id |
字符串 | 对应SQLite文件的数据库名称 |
question |
字符串 | 用户提出的自然语言问题 |
evidence |
字符串或空值 | 解释问题所需的支持信息或定义 |
SQL |
字符串 | 经验证可成功执行的真实SQL查询 |
difficulty |
字符串 | 难度级别:simple、moderate或challenging |
数据获取
新用户
下载完整数据库包: https://drive.google.com/file/d/13VLWIwpw5E3d5DUkMvzw7hvHE67a4XkG/view?usp=sharing
加载数据集
python from datasets import load_dataset dataset = load_dataset("birdsql/bird_sql_dev_20251106") print(dataset["dev_20251106"][0])
基线性能
| 模型 | Dev 1106 |
|---|---|
| claude-sonnet-4.5 | 66.56 |
| gemini-2.5-flash | 65.91 |
| qwen3-coder-480b-a35b | 65.45 |
| claude-sonnet-4 | 64.86 |
| gemini-2.0-flash-001 | 63.62 |
| gpt-5-2025-08-07 | 63.3 |
| Qwen3-30B-A3B-Instruct-2507 | 63.17 |
| Qwen3-235B-A22B-Thinking-2507 | 61.6 |
| Qwen2.5-Coder-32B-Instruct | 60.95 |
| claude-4-5-haiku | 60.69 |
| Llama-3.1-70B-Instruct | 59.39 |
| Qwen2.5-Coder-14B-Instruct | 57.04 |
| Qwen2.5-Coder-7B-Instruct | 49.22 |
| Llama-3.1-8B-Instruct | 36.7 |
引用
bibtex @article{li2024can, title={Can llm already serve as a database interface? a big bench for large-scale database grounded text-to-sqls}, author={Li, Jinyang and Hui, Binyuan and Qu, Ge and Yang, Jiaxi and Li, Binhua and Li, Bowen and Wang, Bailin and Qin, Bowen and Geng, Ruiying and Huo, Nan and others}, journal={Advances in Neural Information Processing Systems}, volume={36}, year={2024} }




