ALeRCE text-to-SQL dataset
收藏资源简介:
ALeRCE文本到SQL数据集由智利大学等研究机构联合构建,旨在为天文数据库的自然语言查询提供基准支持。该数据集包含110对精心设计的自然语言问题与对应SQL查询,数据来源于ALeRCE天文数据库,该数据库整合了兹威基瞬变设施和维拉·鲁宾天文台的观测数据流。数据集的创建过程涉及领域专家手工编写问题与查询对,以确保其覆盖不同复杂度的天文查询场景。该数据集主要应用于评估大型语言模型在天文领域的文本到SQL解析能力,解决天文数据查询的专业门槛问题,推动科学数据库的民主化访问。
The ALeRCE Text-to-SQL dataset was jointly developed by research institutions including the University of Chile, with the goal of providing benchmark support for natural language queries against astronomical databases. This dataset consists of 110 carefully crafted pairs of natural language questions and their corresponding SQL queries, sourced from the ALeRCE astronomical database, which integrates observational data streams from the Zwicky Transient Facility and the Vera C. Rubin Observatory. The dataset construction process involved domain experts manually composing these question-query pairs to ensure coverage of astronomical query scenarios with diverse complexity levels. This dataset is primarily applied to evaluate the text-to-SQL parsing capabilities of large language models in the astronomical domain, addressing the professional threshold barrier for astronomical data querying and promoting democratized access to scientific databases.




