TREECUT
收藏资源简介:
TREECUT是一个合成数据集,由哥伦比亚大学的研究人员创建,旨在生成具有特定结构的数学文字问题。该数据集包含无限数量的无答案数学问题及其对应的有答案版本,通过从一个有答案的问题中移除特定的必要条件来生成无答案的问题。数据集的设计允许精确控制问题的结构组件,如变量数量、问题深度、实体名称的复杂性等,从而为研究大型语言模型在数学推理方面的能力提供了有力的工具。
TREECUT is a synthetic dataset created by researchers at Columbia University, designed to generate mathematical word problems with specific structures. This dataset includes an unlimited number of unanswered mathematical word problems and their corresponding answered variants, where the unanswered problems are generated by removing specific necessary conditions from an existing answered question. The dataset's design allows for precise control over the structural components of the problems, such as the number of variables, problem depth, complexity of entity names, and more, thus serving as a powerful tool for studying the mathematical reasoning capabilities of large language models.




