遇见数据集

SanskritGEC: A Grammatical Error Correction Dataset for Sanskrit (V1.0 + V2.0)

收藏
Zenodo2026-07-05 更新2026-08-01 收录
官方服务:

资源简介:

This deposit contains two datasets for Sanskrit Grammatical Error Correction (GEC), introduced in "Preserving Sanskrit: Dataset Curation and AI-Based Fine-Tuned Model for Grammar Error Correction." SanskritGEC_V1.0 is a manually curated dataset of 855 sentence pairs (correct/erroneous), covering seven grammatical error types, sourced from the grammar exercise textbook Rachanānuvādakaumudī (Dvivedi, 2024). SanskritGEC_V2.0 is a large synthetic dataset of 103,798 sentence pairs, generated using linguistically-inspired error generation methods, designed for fine-tuning transformer-based GEC systems. Both datasets use IAST (International Alphabet of Sanskrit Transliteration) encoding and are intended for training and evaluating grammar error correction, morphological analysis, and NLP tools for morphologically rich, low-resource languages. Access to these datasets is currently restricted; requests are reviewed and approved on a case-by-case basis pending publication of the associated journal article.

提供机构:
Zenodo
创建时间:
2026-07-05
二维码
社区交流群
二维码
科研交流群
商业服务