遇见数据集

RohithMidigudla/gemma-health-medical-sft-balanced

收藏
Hugging Face2026-05-16 更新2026-05-31 收录
官方服务:

资源简介:

Gemma Health Telugu SFT是一个专注于健康和医疗领域的泰卢固语(Telugu)文本生成数据集,用于监督微调(SFT)。数据集包含146,822行训练数据和36,000行测试数据,每行数据以对话格式组织,包括messages字段(采用TRL/Unsloth对话SFT格式)、text字段(作为纯文本序列化聊天的回退),以及用于可追溯性的source、variant、prompt和response字段。该数据集旨在支持自然语言处理任务,特别是在泰卢固语医疗文本生成方面的应用。

Gemma Health Telugu SFT is a Telugu text-generation dataset focused on health and medical domains, designed for supervised fine-tuning (SFT). It includes 146,822 rows for training and 36,000 rows for testing, with each row structured in a conversational format featuring messages (in TRL/Unsloth conversational SFT format), text (as a plain serialized chat text fallback), and traceability fields such as source, variant, prompt, and response. The dataset supports natural language processing tasks, particularly for medical text generation in Telugu.

提供机构:
RohithMidigudla
二维码
社区交流群
二维码
科研交流群
商业服务