遇见数据集

Lu-lab/sintetica-lab-es

收藏
Hugging Face2026-04-28 更新2026-05-03 收录
官方服务:

资源简介:

Sintetica Lab西班牙语样本是一个高质量的合成数据集,用于训练西班牙语大型语言模型。数据集包含多个领域的对话式问答对,如银行、健康、电子商务、法律、教育和人力资源,每个领域提供20个免费样本。数据以JSONL格式提供,适用于文本生成和问答任务,生成过程中使用Ollama模型确保数据隐私。此外,支持用户定制更大规模的数据集。

Sintetica Lab — Spanish Samples is a high-quality synthetic dataset designed for training Spanish large language models (LLMs). It includes conversational question-answer pairs across multiple domains such as banking, healthcare, eCommerce, legal, education, and HR, with 20 free samples per domain. The data is provided in JSONL format, suitable for text-generation and question-answering tasks, and generated using Ollama models to ensure no third-party private data is used. Custom orders for larger datasets are also available.

提供机构:
Lu-lab
二维码
社区交流群
二维码
科研交流群
商业服务