ATH-MaaS/Parrot-dataset
收藏资源简介:
# Description The Parrot dataset is a multilingual, multimodal dataset comprising two parts: the multimodal training datasets sharegpt-4v-ar, sharegpt-4v-pt, sharegpt-4v-ru, sharegpt-4v-tr, and sharegpt-4v-zh, as well as the multimodal evaluation benchmarks MMBench and MMMB. For detailed information about the dataset, please refer to: [Parrot](https://arxiv.org/abs/2406.02539). For the images in the ShareGPT and MMBench datasets, you can refer to the original datasets to obtain them. Due to translation and review processes, the number of data points in other translated languages in mmbench will be fewer than in the original English dataset (each language will have more than 95%). The images in the MMMB dataset have been encoded in base64, so they need to be decoded from base64 for use. # License The dataset is released under CC BY-NC-SA 4.0. The data is released for non-commercial research purposes only. # Declaration Data Sources: - We use data from ShareGPT4v (https://huggingface.co/datasets/Lin-Chen/ShareGPT4V) under Attribution-NonCommercial 4.0 International (https://creativecommons.org/licenses/by-nc/4.0/legalcode.en), and it should abide by the policy of OpenAI (https://openai.com/policies/terms-of-use). - We use data from MME (https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/Evaluation). Copyright belongs to the original dataset owners. - We use data from MMBench (https://github.com/open-compass/MMBench) under Apache License Version 2.0 (https://github.com/open-compass/MMBench/blob/main/LICENSE). Copyright belongs to the original dataset owners. - We use data from ScienceQA (https://github.com/lupantech/ScienceQA?tab=readme-ov-file) under CC BY-NC-SA 4.0 (https://github.com/lupantech/ScienceQA/blob/main/LICENSE-DATA). - We use data from SEED-Bench (https://github.com/AILab-CVC/SEED-Bench?tab=readme-ov-file) under Apache License Version 2.0 (https://github.com/AILab-CVC/SEED-Bench?tab=License-1-ov-file). Copyright belongs to the original dataset owners. Please contact us if you believe any data infringes upon your rights, and we will remove it.
# 数据集概述 Parrot数据集是一款多语言多模态数据集,由两部分构成:多模态训练数据集sharegpt-4v-ar、sharegpt-4v-pt、sharegpt-4v-ru、sharegpt-4v-tr及sharegpt-4v-zh,以及多模态评估基准MMBench与MMMB。 如需了解该数据集的详细信息,请参阅论文《Parrot》(https://arxiv.org/abs/2406.02539)。 ShareGPT与MMBench数据集中的图像可通过其原始数据集获取。由于翻译与审核流程,MMBench中其他翻译语言的样本数量少于原始英文数据集(单语言样本覆盖率均超过95%)。MMMB数据集中的图像已采用base64编码,使用前需进行base64解码。 # 授权协议 本数据集采用CC BY-NC-SA 4.0协议发布,仅可用于非商业性研究用途。 # 声明 数据来源: - 本数据集使用ShareGPT4V(https://huggingface.co/datasets/Lin-Chen/ShareGPT4V)的数据,该数据遵循Attribution-NonCommercial 4.0 International协议(https://creativecommons.org/licenses/by-nc/4.0/legalcode.en),同时需遵守OpenAI(https://openai.com/policies/terms-of-use)的相关条款。 - 本数据集使用MME(https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/Evaluation)的数据,版权归原数据集所有者所有。 - 本数据集使用MMBench(https://github.com/open-compass/MMBench)的数据,该数据遵循Apache License Version 2.0协议(https://github.com/open-compass/MMBench/blob/main/LICENSE),版权归原数据集所有者所有。 - 本数据集使用ScienceQA(https://github.com/lupantech/ScienceQA?tab=readme-ov-file)的数据,该数据遵循CC BY-NC-SA 4.0协议(https://github.com/lupantech/ScienceQA/blob/main/LICENSE-DATA)。 - 本数据集使用SEED-Bench(https://github.com/AILab-CVC/SEED-Bench?tab=readme-ov-file)的数据,该数据遵循Apache License Version 2.0协议(https://github.com/AILab-CVC/SEED-Bench?tab=License-1-ov-file),版权归原数据集所有者所有。 若您认为本数据集包含侵犯您权益的内容,请联系我们,我们将及时移除相关数据。



