遇见数据集

MS-Bench

收藏
DataONE2025-05-14 更新2025-11-01 收录
官方服务:

资源简介:

This is MS-Bench, the first comprehensive benchmark co-developed with archaeologists, comprising 5,076 high-resolution images from 4th to 14th century and 9,982 expert-curated questions across nine sub-tasks aligned with archaeological workflows. Through four prompting strategies, we systematically evaluate 32 LMMs on their: effectiveness robustness cultural contextualization

本数据集为MS-Bench,是首个与考古学家联合研发的综合基准测试集。其包含公元4世纪至14世纪的5076张高分辨率图像,以及9982道由专家精心编撰的问题,涵盖9个贴合考古工作流程的子任务。研究团队通过四种提示策略,系统评估了32款大语言多模态模型(Large Multimodal Model, LMM)在有效性、鲁棒性与文化语境适配能力三个维度的性能表现。

创建时间:
2025-10-29
搜集汇总
数据集介绍
MS-Bench 数据集图片
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务