Lu-lab/sintetica-lab-es
收藏官方服务:
资源简介:
Sintetica Lab西班牙语样本是一个高质量的合成数据集,用于训练西班牙语大型语言模型。数据集包含多个领域的对话式问答对,如银行、健康、电子商务、法律、教育和人力资源,每个领域提供20个免费样本。数据以JSONL格式提供,适用于文本生成和问答任务,生成过程中使用Ollama模型确保数据隐私。此外,支持用户定制更大规模的数据集。
Sintetica Lab — Spanish Samples is a high-quality synthetic dataset designed for training Spanish large language models (LLMs). It includes conversational question-answer pairs across multiple domains such as banking, healthcare, eCommerce, legal, education, and HR, with 20 free samples per domain. The data is provided in JSONL format, suitable for text-generation and question-answering tasks, and generated using Ollama models to ensure no third-party private data is used. Custom orders for larger datasets are also available.
提供机构:
Lu-lab


