SanskritGEC: A Grammatical Error Correction Dataset for Sanskrit (V1.0 + V2.0)
收藏资源简介:
This deposit contains two datasets for Sanskrit Grammatical Error Correction (GEC), introduced in "Preserving Sanskrit: Dataset Curation and AI-Based Fine-Tuned Model for Grammar Error Correction." SanskritGEC_V1.0 is a manually curated dataset of 855 sentence pairs (correct/erroneous), covering seven grammatical error types, sourced from the grammar exercise textbook Rachanānuvādakaumudī (Dvivedi, 2024). SanskritGEC_V2.0 is a large synthetic dataset of 103,798 sentence pairs, generated using linguistically-inspired error generation methods, designed for fine-tuning transformer-based GEC systems. Both datasets use IAST (International Alphabet of Sanskrit Transliteration) encoding and are intended for training and evaluating grammar error correction, morphological analysis, and NLP tools for morphologically rich, low-resource languages. Access to these datasets is currently restricted; requests are reviewed and approved on a case-by-case basis pending publication of the associated journal article.



