MELLA
收藏资源简介:
MELLA是一个多模态、多语言数据集,旨在解决低资源语言环境中多模态大型语言模型(MLLMs)的性能问题。该数据集包含680万个图像-文本对,覆盖阿拉伯语、捷克语、匈牙利语、韩语、俄语、塞尔维亚语、泰语和越南语等八种低资源语言。MELLA数据集的独特之处在于其双源策略,通过收集原始网络HTML的alt-text以及MLLM生成的详细英文图像描述,分别构建了文化知识和语言能力两个子数据集,从而在低资源语言环境中有效地提升了MLLM的语言能力和文化适应性。
MELLA is a multimodal, multilingual dataset designed to address the performance issues of multimodal large language models (MLLMs) in low-resource language contexts. This dataset contains 6.8 million image-text pairs, covering eight low-resource languages including Arabic, Czech, Hungarian, Korean, Russian, Serbian, Thai, and Vietnamese. The uniqueness of the MELLA dataset lies in its dual-source strategy: it constructs two sub-datasets focused on cultural knowledge and linguistic competence respectively by collecting both the alt-text from raw web HTML and detailed English image descriptions generated by MLLMs, thereby effectively enhancing the linguistic proficiency and cultural adaptability of MLLMs in low-resource language environments.




