RAFT for Georgian: A High-Quality Corpus for RAG, Fine-Tuning, and Benchmarking in a Low-Resource Language
收藏资源简介:
Abstract Retrieval-augmented generation (RAG) has shown significant promise in improving factuality and grounding in large language models. However, applying these techniques to low-resource languages like Georgian remains challenging due to the lack of curated, high-quality data. In this project, we present a novel Georgian-language corpus tailored for RAG, built using a RAFT-inspired methodology. We sourced over 2,000 high-quality question-answer pairs from advanced Georgian-language computer science textbooks and academic literature, curated with the support of domain experts from local CS faculties.The dataset underwent multiple stages of filtering and quality assurance, including manual verification by trained human annotators to ensure factual consistency, relevance, and diversity.This corpus represents one of the first structured, RAG-friendly datasets for the Georgian language and is designed to facilitate grounded reasoning, improved retrieval, and better downstream performance of Georgian LLMs.It can be used both as a high-quality dataset for fine-tuning Georgian models and as a benchmark for evaluating retrieval and reasoning performance in Georgian.Our results demonstrate that targeted curation, even at smaller scales, can significantly improve outcomes in specialized domains and under-resourced languages. Acknowledgements The work was supported by the European Union Horizon Europe grant project GAIN.



