BREAK
收藏资源简介:
BREAK数据集是由特拉维夫大学等机构创建的,包含83,978个问题及其对应的QDMR(Question Decomposition Meaning Representation)分解。该数据集通过众包方式收集,涵盖了从10个不同数据集和三种信息源(数据库、文本、图像)中随机抽样的问题。BREAK数据集的创建旨在推动问题理解模型的发展,特别是在需要多步骤推理或大量数据不可用的情况下。数据集的应用领域包括开放领域问答和语义解析,旨在通过问题分解提高这些任务的性能和泛化能力。
The BREAK dataset was developed by Tel Aviv University and other institutions, containing 83,978 questions and their corresponding QDMR (Question Decomposition Meaning Representation) decompositions. Collected through crowdsourcing, this dataset covers questions randomly sampled from 10 distinct datasets and three types of information sources: databases, text, and images. The BREAK dataset is designed to promote the advancement of question understanding models, especially in scenarios requiring multi-step reasoning or where large-scale data is unavailable. Its application domains include open-domain question answering and semantic parsing, with the aim of enhancing the performance and generalization capability of these tasks via question decomposition.



