Urdu Multi-Domain Benchmark: Nastaliq and Roman Urdu Classification Domains for Domain-Robustness Evaluation
收藏资源简介:
A benchmark for domain robustness of Urdu text classifiers: 26 datasets in Nastaliq and Roman Urdu, grouped into 7 transfer matrices over four task types (sentiment and polarity, abusive language, fake news, question-answer pair validation); 188,309 labelled examples. Version 1.1 supersedes v1.0. Domains were selected by explicit inclusion criteria (no duplicated sources, no labels predictable from the origin of an item, at least 100 test items), and every split is free of exact and near-duplicate train-test overlap (cluster-aware splits for LLM-generated domains, question-grouped splits for QA). Licences. Each domain keeps the licence of its source (see LICENSE.md and domains.csv). Four public corpora whose authors state no licence (Urdu Sentiment Corpus, ISE-Hate, Bend the Truth, Ax-to-Grind Urdu) are distributed as manifests plus rebuild.py, which downloads the authors' own files. Generated domains were produced with xAI Grok and are attributed to it. Our own contributions (splits, manifests, audits, curation) are released under CC BY-NC 4.0. Content warning: the abusive-language domains contain offensive, hateful and sectarian text. Code, audits and model predictions: https://github.com/AbdullahPatti/Urdu-Multi-Domain-Script-Research



