MS-Bench
收藏数据链接:
官方服务:
资源简介:
This is MS-Bench, the first comprehensive benchmark co-developed with archaeologists, comprising 5,076 high-resolution images from 4th to 14th century and 9,982 expert-curated questions across nine sub-tasks aligned with archaeological workflows. Through four prompting strategies, we systematically evaluate 32 LMMs on their: effectiveness robustness cultural contextualization
本数据集为MS-Bench,是首个与考古学家联合研发的综合基准测试集。其包含公元4世纪至14世纪的5076张高分辨率图像,以及9982道由专家精心编撰的问题,涵盖9个贴合考古工作流程的子任务。研究团队通过四种提示策略,系统评估了32款大语言多模态模型(Large Multimodal Model, LMM)在有效性、鲁棒性与文化语境适配能力三个维度的性能表现。
创建时间:
2025-10-29
搜集汇总
数据集介绍

以上内容由遇见数据集搜集并总结生成



