Awesome-Text2SQL-Dataset
收藏资源简介:
Awesome-Text2SQL-Dataset 是一个专门为Text-to-SQL任务策划的数据集集合,该任务涉及将自然语言问题转换为SQL查询。该合集旨在为研究人员、开发者和从业者提供全面的数据集列表,以支持和评估Text2SQL模型,涵盖模式链接、复杂SQL生成、跨领域泛化和对话查询生成等领域。它包含最新的数据集资源(如2026年和2025年的条目),并持续更新以服务社区。
Awesome-Text2SQL-Dataset is a curated dataset collection specifically designed for the Text-to-SQL task, which involves converting natural language questions into SQL queries. This collection aims to provide researchers, developers and practitioners with a comprehensive list of datasets to support and evaluate Text2SQL models, covering fields such as schema linking, complex SQL generation, cross-domain generalization and conversational query generation. It includes the latest dataset resources (such as entries from 2025 and 2026) and is continuously updated to serve the community.
Awesome-Text2SQL-Dataset 数据集详情总结
该仓库是一个精心整理的 Text-to-SQL 数据集集合,旨在为将自然语言问题转换为 SQL 查询的任务提供支持,涵盖模式链接、复杂 SQL 生成、跨领域泛化及对话式查询生成等多个方向。
最新数据集(2026年)
-
SQL Injection Dataset(2026/01)
-
BIRDTurk(2026/01)
2025年数据集
-
DAComp(2025/12)
-
DP-Bench(2025/12)
-
DLBench(2025/11)
-
BibSQL(2025/11)
-
DySQL-Bench(2025/11)
-
DeKeyNLU(2025/11)
-
GeoSQL-Bench(2025/11)
-
DBASQL(2025/10)
-
CORGI(2025/10)
-
Payment-SQL(2025/09)
-
Arabic WikiTableQA(2025/09)
-
LLMSQL(2025/09)
-
text2SQL4PM(2025/08)
-
REEF(2025/08)
-
CogniSQL(2025/07)
-
SQLStorm(2025/07)
-
BIRD-Critic(2025/06)
-
BiomedSQL(2025/05)
-
LogicCat(2025/05)
-
TINYSQL(2025/03)
-
NL2SQL-Bugs(2025/03)
-
OmniSQL / SynSQL-2.5M(2025/03)
其他重要数据集
-
WikiSQL(2017/09)
-
Spider 1.0(2018/09)
-
SParC(2019/06)
-
CSpider(2019/09)
-
CoSQL(2019/09)
-
KaggleDBQA(2021/06)
-
Spider-Syn(2021/06)
-
SEDE(2021/06)
-
CHASE(2021/08)
-
Spider-DK(2021/09)



