SEED-Bench
收藏资源简介:
# SEED-Bench Card ## Benchmark details **Benchmark type:** SEED-Bench is a large-scale benchmark to evaluate Multimodal Large Language Models (MLLMs). It consists of 19K multiple choice questions with accurate human annotations, which covers 12 evaluation dimensions including the comprehension of both the image and video modality. **Benchmark date:** SEED-Bench was collected in July 2023. **Paper or resources for more information:** https://github.com/AILab-CVC/SEED-Bench **License:** Attribution-NonCommercial 4.0 International. It should abide by the policy of OpenAI: https://openai.com/policies/terms-of-use. For the images of SEED-Bench, we use the data from Conceptual Captions Dataset (https://ai.google.com/research/ConceptualCaptions/) following its license (https://github.com/google-research-datasets/conceptual-captions/blob/master/LICENSE). Tencent does not hold the copyright for these images and the copyright belongs to the original owner of Conceptual Captions Dataset. For the videos of SEED-Bench, we use tha data from Something-Something v2 (https://developer.qualcomm.com/software/ai-datasets/something-something), Epic-kitchen 100 (https://epic-kitchens.github.io/2023) and Breakfast (https://serre-lab.clps.brown.edu/resource/breakfast-actions-dataset/). We only provide the video name. Please download them in their official websites. **Where to send questions or comments about the benchmark:** https://github.com/AILab-CVC/SEED-Bench/issues ## Intended use **Primary intended uses:** The primary use of SEED-Bench is evaluate Multimodal Large Language Models on spatial and temporal understanding. **Primary intended users:** The primary intended users of the Benchmark are researchers and hobbyists in computer vision, natural language processing, machine learning, and artificial intelligence.
<p align="center" width="100%"> <img src="https://i.postimg.cc/g0QRgMVv/WX20240228-113337-2x.png" width="100%" height="80%"> </p> # 大规模多模态模型评测套件 > 借助`lmms-eval`加速大规模多模态模型(Large-scale Multi-modality Models, LMMs)的研发 🏠 [主页](https://lmms-lab.github.io/) | 📚 [文档](docs/README.md) | 🤗 [Huggingface 数据集仓库](https://huggingface.co/lmms-lab) # 本数据集 本数据集是[SEED-Bench](https://github.com/AILab-CVC/SEED-Bench)的格式化版本,可集成至我们的`lmms-eval`流水线,实现大规模多模态模型的一键评测。 @article{li2023seed, title={SEED-Bench:面向生成式理解能力的多模态大语言模型评测基准}, author={Li, Bohao and Wang, Rui and Wang, Guangzhi and Ge, Yuying and Ge, Yixiao and Shan, Ying}, journal={arXiv预印本 arXiv:2307.16125}, year={2023} }




