Urdu–Shahmukhi Punjabi Bilingual Lemmatization Dataset and Benchmark
收藏资源简介:
This dataset provides a bilingual collection of word–lemma pairs for Urdu and Shahmukhi Punjabi, two morphologically rich Indo-Aryan languages. The dataset consists of pairs of inflected word forms and their corresponding lemmas, represented at the character level. It is designed to support sequence-to-sequence modeling approaches for lemmatization and morphological analysis. The resource combines:- An Urdu morphological dataset derived from an existing lexical resource- A Shahmukhi Punjabi dataset curated and annotated for this study Each entry in the dataset contains:- Word: an inflected surface form- Lemma: the corresponding base form The dataset is suitable for:- character-level lemmatization tasks- morphological modeling in low-resource languages- bilingual and cross-lingual NLP experiments The dataset is provided in a unified tabular format. Users can create their own train, validation, and test splits depending on their experimental setup. This dataset is released to support reproducible research in low-resource NLP. Documentation is provided describing data format, preprocessing steps, and usage instructions. Note: The Urdu portion of the dataset is derived from an existing morphological resource. Any licensing considerations associated with the original data source are described in the accompanying documentation.



