Rosetta Stone–Match-Up paired puzzles corpus
收藏资源简介:
该数据集由纽约市立大学等机构创建,是一个专门针对语言学谜题的配对语料库,核心包含Rosetta Stone和Match-Up两种常见谜题格式的对应版本。数据集内容源于英国语言学奥林匹克竞赛(UKLO)发布的原始Rosetta Stone谜题及其官方解答,通过系统化转换流程生成了对应的Match-Up版本,从而形成了结构化的谜题对。其创建过程涉及对原始谜题陈述、问题及答案的提取,并遵循特定规则进行格式转换与配对。该数据集主要应用于评估人类与大语言模型在语言学推理任务上的表现,旨在探究不同谜题格式是否代表本质相同的底层结构,并为高效的谜题生成方法提供基准测试资源。
This dataset, developed by institutions including the City University of New York, is a paired corpus dedicated to linguistic puzzles, with core contents consisting of matched versions of two prevalent puzzle formats: Rosetta Stone and Match-Up. The dataset’s content is derived from the original Rosetta Stone puzzles and their official solutions published by the United Kingdom Linguistics Olympiad (UKLO). Corresponding Match-Up versions were generated through a systematic conversion pipeline, yielding structured puzzle pairs. Its construction process entails extracting the original puzzle descriptions, questions, and answers, followed by format conversion and pairing following specific guidelines. This dataset is primarily utilized to assess the performance of both humans and large language models (LLMs) on linguistic reasoning tasks. It seeks to investigate whether distinct puzzle formats embody fundamentally identical underlying structures, and serves as a benchmark resource for efficient puzzle generation methodologies.




