JSONSchemaBench
收藏资源简介:
JSONSchemaBench是一个由瑞士洛桑联邦理工学院和微软等机构共同创建的数据集,旨在评估语言模型在生成结构化输出时的性能。该数据集包含10,000个真实世界的JSON模式,涵盖了从简单到复杂的多种约束类型,适用于函数签名、服务API和系统配置等领域。数据集的创建过程包括从公开的GitHub仓库、JSON Schema测试套件等来源收集数据,并经过标准化处理以确保一致性。该数据集的应用领域主要集中在结构化生成任务中,旨在解决语言模型在生成符合预定义格式和约束的输出时的挑战。
JSONSchemaBench is a dataset co-created by institutions including École Polytechnique Fédérale de Lausanne (EPFL) of Switzerland and Microsoft, aiming to evaluate the performance of language models when generating structured outputs. This dataset contains 10,000 real-world JSON schemas, covering a wide range of constraint types from simple to complex, and is applicable to domains such as function signatures, service APIs, and system configurations. The creation process of this dataset involves collecting data from sources including public GitHub repositories and JSON Schema test suites, followed by standardization processing to ensure consistency. The application scenarios of this dataset mainly focus on structured generation tasks, aiming to address the challenges faced by language models when generating outputs that comply with predefined formats and constraints.

- 1Generating Structured Outputs from Language Models: Benchmark and Studies瑞士洛桑联邦理工学院, 微软, JSON Schema · 2025年



