Shruti0603/birdbench
收藏资源简介:
BirdBench数据集是一个用于自然语言处理(NLP)任务的大规模英语数据集,专注于文本到SQL转换(text2sql)、问答、表格问答和文本生成。它支持多种任务类别,包括问题回答、表格问题回答、文本生成和文本到文本生成,并设计用于与大型语言模型(如LLaMA)集成,涉及数据库相关应用。数据规模估计在1亿到10亿个元素之间,适用于训练和评估AI模型在SQL查询生成和其他文本处理方面的能力。
BirdBench is a large-scale English dataset dedicated to natural language processing (NLP) tasks, focusing on text-to-SQL conversion (text2sql), question answering, table question answering, and text generation. It supports a diverse range of task categories, including question answering, table question answering, text generation, and text-to-text generation, and is designed for integration with large language models (e.g., LLaMA) for database-related applications. Its scale is estimated to range from 100 million to 1 billion data points, and it is suitable for training and evaluating the capabilities of AI models in SQL query generation and other text processing tasks.



