IslamicRAG-7M: A retrieval-ready multilingual embedding corpus of Quran, Hadith, video lectures, and scholarly literature
收藏资源简介:
We present IslamicRAG-7M, a retrieval-ready multilingual embedding corpus that spans four heterogeneous sources of Islamic textual content: the Quran (6,236 verses with Arabic, English, and Bangla parallel material), the six major Sunni hadith collections together with Muwatta Malik (34,552 hadith with Arabic and English text), transcribed YouTube lectures from three prominent Islamic scholars (31,262 timestamped chunks in Bangla and English), and 31,118 books of classical and contemporary Islamic scholarly literature drawn from Al-Maktaba Al-Shamela and two curated Deobandi-Hanafi collections (7,720,380 OCR-extracted pages). Every record carries a 768-dimensional dense embedding computed with Google's gemini-embedding-001 model using the RETRIEVAL_DOCUMENT task type, yielding a unified 7.79-million-vector semantic index across Arabic, English, Bangla, and Urdu. Each PDF book is additionally enriched with LLM-extracted bibliographic metadata (author, publisher, year, language, description) and a two-level canonical subject classification over a fixed 22-category Islamic-studies taxonomy. The dataset is released in Apache Parquet format with strict source attribution, positioning it as a reproducible foundation for retrieval-augmented generation, semantic search, cross-lingual knowledge alignment, and computational studies of Islamic scholarship.



