IslamicTurathBench (ISTB): a benchmark for the Islamic Scholarly Tradition in Arabic
收藏资源简介:
IslamicTurathBench (ISTB) is an expert-authored Arabic benchmark dataset for evaluating large language models on the Islamic scholarly tradition (turāth). It contains 3,465 question–answer items across seven fields of Islamic studies, stratified by difficulty level and grounded in 35 named source works used in Islamic scholarly learning. The benchmark measures whether computational systems can engage with source-based, disciplinary, and interpretive knowledge from the Islamic scholarly tradition, rather than only demonstrating general Arabic fluency or answering generic religious questions. Arabic is the language of the dataset; the subject domain is Islamic scholarship. This dataset is described in the following preprint: Gaben, S., Sbahi, H., Rashwani, S., Bouchekif, A., Al-Khatib, M., Mohamed, E., Eltanbouly, S., & Ghaly, M. (2026). IslamicTurathBench: A Multi-Task, Multi-Discipline Benchmark for Evaluating Large Language Models on the Islamic Scholarly Tradition (turath). arXiv preprint arXiv:2608.04703. https://arxiv.org/abs/2608.04703 The manuscript has also been submitted to Scientific Data and is currently under consideration. Task types ISTB includes three task formats: MCQ — multiple-choice question answering: 2,276 items COMP — passage-based open-ended comprehension: 417 items KNOW — closed-book open-ended question answering: 772 items Disciplines The dataset covers seven fields of Islamic studies, with near-balanced coverage across disciplines: Qurʾanic sciences (ʿulūm al-Qurʾān) Hadith sciences (ʿulūm al-Ḥadīth) Islamic theology (ʿaqīda) Sufism (taṣawwuf) Principles of jurisprudence (uṣūl al-fiqh) Jurisprudence (fiqh) Prophetic biography (sīra) Items are stratified across three difficulty levels: Beginner, Intermediate, and Advanced. These levels reflect increasing scholarly depth, from introductory works to standard disciplinary works, commentaries, advanced treatises, major commentaries, and major disciplinary works. Jurisprudence items are bounded by the selected source works included in the dataset and do not claim exhaustive coverage of all legal schools or subtraditions. Authorship and source grounding Questions and gold answers were developed and reviewed by domain experts, including doctoral specialists in Islamic Studies; they were not web-scraped. COMP items include short Arabic excerpts, not full books, from the reference works named in each row and in source_links.csv. Human reference panel A scholarly human reference panel is included for comparative benchmarking. It consists of a stratified 148-question subset covering all seven disciplines, all three difficulty levels, and all three task formats, together with de-identified aggregate panel scores. The human reference layer is intended as a generalist scholarly reference baseline, not as a specialist ceiling. Individual evaluator responses and personally identifiable information are not included. Deposit contents The deposit provides: Three task collections in both JSON and CSV format, UTF-8 encoded Bibliographic metadata for 35 reference works: source_links.csv Corpus statistics: statistics/unified_data_stats.json Human reference panel subset and aggregate scores: human_panel/ A minimal Python loader: load_benchmark.py A single manifest: MANIFEST.json file inventory per-task column schema row counts SHA-256 checksums A CC BY 4.0 licence No proprietary format or tooling is required. Companion GitHub repository A companion GitHub repository illustrates how the dataset was used in our model benchmarking experiments and how researchers can load and evaluate model performance on ISTB: https://github.com/GabenS99/IslamicTurathBench_Evaluation



