Archit00/jamendo-qa-mirror
收藏资源简介:
Jamendo-QA是一个大规模音乐问答数据集,专为音乐相关问答研究而设计。它基于Jamendo Music音乐集合构建,支持音乐知识问答、音频-文本多模态学习以及音乐信息检索等任务。数据集包含7,335个唯一的音乐曲目,总计约400小时的音频时长,覆盖超过35种音乐流派和7,000多位独特艺术家。数据格式包括Parquet文件(其中嵌入了音频字节)和JSON文件,提供丰富的元数据如艺术家、流派、速度、性别、语言、歌词等,以及两个版本的问答对:qa_v1包含29,340个基于基本元数据的问答对,qa_v2包含58,680个带有详细音乐分析和描述的问答对,总计88,020个问答对。数据集仅用于研究目的,适用于问答、检索增强生成和多模态音乐理解等应用场景。
Jamendo-QA is a large-scale dataset designed for music-related question answering research. It is built upon the Jamendo Music collection and supports research in music knowledge QA, audio-text multimodal learning, and music information retrieval. The dataset contains 7,335 unique music tracks, totaling approximately 400 hours of audio, covering over 35 music genres and more than 7,000 unique artists. The data format includes Parquet files with embedded audio bytes and JSON files, providing rich metadata such as artist, genre, speed, gender, language, lyrics, and two versions of QA pairs: qa_v1 with 29,340 basic metadata-based QA pairs, and qa_v2 with 58,680 detailed music analysis and captioning QA pairs, totaling 88,020 QA pairs. The dataset is available for research-only purposes and is suitable for tasks like question answering, retrieval-augmented generation, and multimodal music understanding.




