JetBrains-Research/cwm-benchmarks-dl4c-generations
收藏资源简介:
该数据集名为CWM Benchmarks — DL4C Generations,包含从v05_clean评估运行中获取的每个模型的原始生成结果,这些结果针对CWM-benchmarks DL4C任务套件进行了评分。每个分割(split)对应一个特定模型,并包含该模型在相同一组(instance, test, side)样本上的435个响应。数据集涵盖了多个模型,如CWM、MiniMax系列、Qwen系列、Claude系列、GPT系列等,总计18个分割。数据以Parquet格式存储,每个样本包括sample_id、task、instance_id、test_nodeid、side、model、duration_s、prompt_tokens、completion_tokens、raw_response、parsed_response、prediction、ground_truth、metric、metric_args、metrics和error等字段。嵌套对象以JSON编码字符串存储,以保持Parquet模式扁平。数据集可用于评估模型在DL4C任务上的性能。
This dataset is named CWM Benchmarks — DL4C Generations. It contains raw generation results of each model obtained from the v05_clean evaluation run, with all results scored against the CWM-benchmarks DL4C task suite. Each split corresponds to a specific model, and contains 435 responses from that model on the same set of (instance, test, side) samples. The dataset covers multiple models, including CWM, MiniMax series, Qwen series, Claude series, GPT series and others, totaling 18 splits. The data is stored in Parquet format, with each sample containing fields such as sample_id, task, instance_id, test_nodeid, side, model, duration_s, prompt_tokens, completion_tokens, raw_response, parsed_response, prediction, ground_truth, metric, metric_args, metrics and error. Nested objects are stored as JSON-encoded strings to maintain a flat Parquet schema. This dataset can be used to evaluate model performance on DL4C tasks.




