遇见数据集

lius-cc/daoism-rag-sft-v1

收藏
Hugging Face2026-05-14 更新2026-05-31 收录
官方服务:

资源简介:

这是一个道教相关的检索增强生成(RAG)感知的监督微调(SFT)数据集,版本为v1。每个数据对都包含检索到的上下文(两份道教文献),旨在训练大型语言模型(LLM)学会基于上下文回答问题。数据格式包括instruction(包含相关道教文献和问题)、input(通常为空)和output(回答)。数据集规模为:训练集344455条,验证集18129条,总计362584条。该数据集主要用于训练LLM以适用于RAG-SaaS场景,在部署到客户自有道教资料库的生产工作流中,相比纯SFT模型能提高30%以上的准确率。

This is a Taoism-related Retrieval-Augmented Generation (RAG)-aware Supervised Fine-Tuning (SFT) dataset, version v1. Each data pair includes retrieved context (two Taoist documents) and is designed to train large language models (LLMs) to answer questions based on context. The data format consists of instruction (containing relevant Taoist documents and questions), input (typically empty), and output (answers). The dataset scale is: 344,455 training samples, 18,129 validation samples, totaling 362,584 samples. It is primarily used for training LLMs in RAG-SaaS scenarios, and when deployed in production workflows with customer-provided Taoist databases, it can improve accuracy by over 30% compared to pure SFT models.

提供机构:
lius-cc
二维码
社区交流群
二维码
科研交流群
商业服务