WASIL
收藏资源简介:
# WASIL Dataset **WASIL** is a dataset of in-the-wild Arabic spoken interactions with an LLM-based assistant. The dataset contains ∼9K turns (9,304 turns from 93 users), spans multiple dialects and countries, and includes explicit user feedback on assistant responses, including like or dislike signals and scalar scores. ## Dataset Overview WASIL contains **9,304 spoken Arabic prompts** from **93 users** interacting with an ASR → LLM assistant. Each interaction includes a **like/dislike** reaction. Disliked responses were further labeled with one or more feedback categories: failed to follow instructions, lacked factual accuracy, displeased with style/format, avoided answering, not aligned with Arabic or Islamic culture, disturbing content, religiously incorrect, grammatical errors, too brief, or too long. ## Released Data ### 1. Test Set (1,416 prompts) A randomly selected subset from the full dataset with additional manual annotations: - **Language/Dialect Labels** — Human-annotated classification covering MSA and four major dialects (Egyptian, Syrian, Algerian and Sudanese) - **Answerability Label** — Human-annotated intrinsic answerability label (answerable, ambiguous/needs-clarification, unsupported, or not-a-request/noise) as a way to separate turns that are answerable from those that require clarification or are out-of-scope - **ASR Hypotheses** — Automatic transcriptions from Fanar and Gemini ASR systems - **Post-edited Transcription** — Human-corrected gold transcripts - **MSA Translation** — LLM-generated Modern Standard Arabic translations for dialectal prompts, manually post-edited by humans ### 2. Feedback Set (988 prompts) A subset from the full dataset containing prompts where users provided explicit feedback on assistant responses. > **Note:** There may be overlap between the Test Set and Feedback Set. ### Data Structure Both released splits share the following structure in the dataset viewer. The `audio` field is generated from the split metadata and points to the corresponding WAV file: ```json { "prompt_id": "<unique identifier>", "prompt_wav_path": "<path to audio file>", "audio": "<playable audio feature>", "prompt_dialect": "<dialect code: EG, SY, DZ, SD, MSA>", "prompt_answerability_label": "<ANSWERABLE_CLEAR | AMBIGUOUS_NEEDS_CLARIFICATION | OUT_OF_DOMAIN_UNSUPPORTED | NOT_A_REQUEST_BACKCHANNEL_NOISE>", "prompt_transcription": { "asr_transcription_fanar": "<Fanar ASR output>", "asr_transcription_gemini": "<Gemini ASR output>", "gold_transcription": "<human-corrected transcription>", "gold_transcription_msa_translation": "<MSA translation of dialectal prompt>" }, "response_id": "<response identifier>", "response": "<assistant's generated response>", "reaction": "<like | dislike>", "feedback": "<user's written feedback text>", "feedback_category": "<category of the feedback>", "feedback_score": "<numeric score>" } ``` ## Citation If you use this dataset, please cite our paper: ```bibtex @article{ali2026wasilinthewildarabicspoken, title = {{WASIL}: In-the-Wild Arabic Spoken Interactions with LLMs}, author = {Ali, Zien Sheikh and Mubarak, Hamdy and Jung, Soon-Gyo and Bhatti, Hunzalah Hassan and Alam, Firoj and Chowdhury, Shammur Absar}, journal = {arXiv preprint arXiv:2605.16364}, year = {2026}, archivePrefix = {arXiv}, eprint = {2605.16364}, primaryClass = {cs.SD}, url = {https://arxiv.org/abs/2605.16364} } ``` ## License This dataset is released under the CC BY-NC-SA 4.0 license.



