Bloom Library
收藏资源简介:
Bloom Library是一个包含多模态和多语言数据集的平台,旨在支持语言建模、图像描述、视觉叙事和语音合成/识别等多种下游任务。该数据集涵盖了32个语系的363种语言,是目前最全面的多语言数据集之一。数据集的内容包括书籍、图像和音频记录,其中许多书籍包含与文本对齐的图像,以及超过1600本书的音频记录。数据集的创建过程涉及与当地语言社区的合作,使用开源软件Bloom进行书籍创作、音频录制和翻译。Bloom Library的应用领域广泛,旨在解决全球语言资源不平等的问题,特别是在低资源语言的NLP研究中建立基准。
Bloom Library is a multimodal and multilingual dataset platform designed to support a wide range of downstream tasks, such as language modeling, image captioning, visual storytelling, and speech synthesis/recognition. It covers 363 languages across 32 language families, making it one of the most comprehensive multilingual datasets currently available. The dataset includes books, images, and audio recordings: many of the books contain images aligned with their corresponding text, and there are audio recordings for over 1,600 books. The development of this dataset involved collaboration with local linguistic communities, and utilized the open-source software Bloom for book creation, audio recording, and translation. Bloom Library has broad application scenarios, aiming to address global disparities in language resources and establish benchmarks for NLP research focused on low-resource languages.




