遇见数据集

socb-anon-2441/socb

收藏
Hugging Face2026-05-14 更新2026-05-31 收录
官方服务:

资源简介:

SOCB(结构化输出一致性基准)是一个用于评估大型语言模型(LLM)在生成结构化输出(如JSON和工具调用)时一致性的数据集。该基准包含多个子集:人类标注的争议案例(74条记录,来自Toucan工具调用任务,由三位标注者标注,Fleiss kappa为0.759)、均匀随机样本(100条记录,用于非争议选择的一致性验证)、合成校准数据(2,957条记录,覆盖语义、表达、模式和顺序四种扰动类别,用于度量标准校准)、外部标签排名三元组(来自BFCL的500条和ShareGPT的429条记录,用于度量排名评估)、SDK强制执行消融输出(13,680条原始生成记录,研究模式对一致性的影响)以及原始生成样本(118,910条记录,代表完整数据集的子集)。数据集总规模约210万条生成记录,涵盖Toucan工具调用和ShareGPT自由格式JSON两种任务,涉及18个模型、11个温度和10次独立运行。数据集旨在支持模型一致性比较、新相似性度量评估和任务内反事实研究,采用CC BY-NC 4.0许可证,仅限非商业使用。

SOCB (Structured Output Consistency Benchmark) is a dataset for evaluating the consistency of Large Language Models (LLMs) when generating structured outputs such as JSON and tool calls. The benchmark includes multiple subsets: human-annotated controversial cases (74 records from the Toucan tool call task, annotated by three annotators with a Fleiss' kappa score of 0.759), uniform random samples (100 records for consistency validation of non-controversial selections), synthetic calibration data (2,957 records covering four perturbation categories: semantic, expressive, schema, and sequential, used for standard calibration measurement), external label ranking triples (500 records from BFCL and 429 records from ShareGPT for ranking evaluation measurement), SDK-enforced ablation outputs (13,680 original generated records for studying the impact of schema on consistency), and raw generated samples (118,910 records representing a subset of the full dataset). The total scale of the dataset is approximately 2.1 million generated records, covering two tasks: Toucan tool call and ShareGPT free-form JSON, involving 18 models, 11 temperature settings, and 10 independent runs. The dataset aims to support model consistency comparison, evaluation of novel similarity metrics, and intra-task counterfactual research, and is licensed under CC BY-NC 4.0 for non-commercial use only.

提供机构:
socb-anon-2441
二维码
社区交流群
二维码
科研交流群
商业服务