遇见数据集

RAFT for Georgian: A High-Quality Corpus for RAG, Fine-Tuning, and Benchmarking in a Low-Resource Language

收藏
Zenodo2025-08-31 更新2026-05-26 收录
官方服务:

资源简介:

Abstract Retrieval-augmented generation (RAG) has shown significant promise in improving factuality and grounding in large language models. However, applying these techniques to low-resource languages like Georgian remains challenging due to the lack of curated, high-quality data. In this project, we present a novel Georgian-language corpus tailored for RAG, built using a RAFT-inspired methodology. We sourced over 2,000 high-quality question-answer pairs from advanced Georgian-language computer science textbooks and academic literature, curated with the support of domain experts from local CS faculties.The dataset underwent multiple stages of filtering and quality assurance, including manual verification by trained human annotators to ensure factual consistency, relevance, and diversity.This corpus represents one of the first structured, RAG-friendly datasets for the Georgian language and is designed to facilitate grounded reasoning, improved retrieval, and better downstream performance of Georgian LLMs.It can be used both as a high-quality dataset for fine-tuning Georgian models and as a benchmark for evaluating retrieval and reasoning performance in Georgian.Our results demonstrate that targeted curation, even at smaller scales, can significantly improve outcomes in specialized domains and under-resourced languages. Acknowledgements The work was supported by the European Union Horizon Europe grant project GAIN.

提供机构:
Zenodo
创建时间:
2025-08-31
二维码
社区交流群
二维码
科研交流群
商业服务