遇见数据集

QSOL-IMC LLM Benchmark v1 — Grok 4.1 vs Claude Sonnet 4.5

收藏
Zenodo2026-02-18 更新2026-05-26 收录
官方服务:

资源简介:

This dataset presents the Phase 1 results of the QSOL-IMC Large Language Model Benchmark, comparing Grok 4.1 and Claude Sonnet 4.5 across 18 carefully designed tasks covering reasoning, mathematics, algorithmic competency, truthfulness, hallucination resistance, and creative constraint satisfaction. Unlike small-sample or anecdotal model comparisons, this benchmark uses domain-balanced prompts and direct side-by-side evaluation of full outputs. Results include a structured CSV table ("Benchmark_Table_v1.csv") summarizing performance metrics for each subtask, with correctness scores, depth-of-reasoning assessments, constraint satisfaction checks, and hallucination integrity evaluations. The benchmark highlights that both models perform strongly in fundamental reasoning and coding tasks, while Grok 4.1 demonstrates greater depth in numerical analysis, algorithmic explanation, and historically sourced fact-checking. This dataset serves as a reproducible baseline for ongoing QSOL-IMC LLM comparative research, with future phases expanding to long-context retention, multi-hop reasoning, adversarial robustness, and domain-specific scientific workloads.

提供机构:
Zenodo
创建时间:
2025-11-24
二维码
社区交流群
二维码
科研交流群
商业服务