CogniSQL Reasoning Traces & Positive Sample Corpus
收藏资源简介:
CogniSQL Reasoning Traces数据集包含5,024个推理轨迹,每个轨迹具有不同的上下文长度,用于研究和训练高效的文本到SQL生成模型。Positive Sample Corpus数据集包含36,356个弱监督查询的正样本语料库,每个查询都标注了六条语义上不同的推理路径。这些数据集由Dell Technologies的研究人员创建,旨在支持可扩展和可解释的文本到SQL生成研究。数据集的大小和多样性为模型训练提供了丰富的资源,同时,数据集的创建过程考虑了执行正确性和格式标签合规性,以确保生成的SQL查询的准确性和可执行性。这些数据集的应用领域主要在于自然语言处理和结构化数据访问,旨在解决将自然语言问题转换为可执行的SQL查询的挑战。
The CogniSQL Reasoning Traces dataset contains 5,024 reasoning traces with varying context lengths, designed for researching and training efficient text-to-SQL generation models. The Positive Sample Corpus dataset comprises 36,356 positive sample corpora for weakly supervised queries, where each query is annotated with six semantically distinct reasoning paths. These datasets were created by researchers at Dell Technologies to support scalable and interpretable text-to-SQL generation research. The size and diversity of the datasets offer abundant resources for model training, and their development process takes execution correctness and format label compliance into consideration to ensure the accuracy and executability of the generated SQL queries. The primary application fields of these datasets cover natural language processing and structured data access, with the goal of addressing the challenge of converting natural language questions into executable SQL queries.




