SINAI/ALIA-es-legal-administrative-synthetic-instructions
收藏资源简介:
ALIA西班牙法律与行政合成指令语料库是一个西班牙语合成指令调优资源,由ALIA项目使用Magpie方法创建。它旨在通过受控格式和大规模监督来训练和评估语言模型在法律和行政任务中的表现。数据集包含764,180个实例和534,721,446个令牌,涵盖16种任务模态,包括问题、指令、多项选择、真/假等,并区分有无上下文和有无理由。该语料库包含合成生成的法律-行政指令-响应对,通过指令模型生成,并经过多步骤质量管道筛选。生成过程遵循Magpie方法,模型首先生成用户侧查询(问题或指令),然后生成相应的答案。语料库包括一般提示(无上下文)和基于西班牙法律和行政文档构建的上下文条件提示,还包含结构化评估友好格式(如多项选择和真/假变体),并明确控制理由。数据集适用于西班牙语法律大语言模型的指令调优、法律和行政问答、推理和受控响应格式的评估,以及领域适应和RAG导向训练的合成监督。
The ALIA Spanish Legal and Administrative Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in legal and administrative tasks with controlled formats and large-scale supervision. It contains 764,180 instances, 534,721,446 tokens, and 16 task modalities (questions, instructions, multiple-choice, true/false; with and without context; with and without justification). This dataset contains synthetic legal-administrative instruction-response pairs generated with an instruction model and curated through a multi-step quality pipeline. The generation process follows Magpie, where the model first produces a user-side query (question or instruction) and then generates the corresponding answer. The corpus includes both general prompts (without context) and context-conditioned prompts built from Spanish legal and administrative documents. It also includes structured evaluation-friendly formats such as multiple-choice and true/false variants, with explicit control over justification. The dataset is intended for instruction tuning of legal LLMs in Spanish, legal and administrative question answering, evaluation of reasoning and controlled response formats, and synthetic supervision for domain adaptation and RAG-oriented training.




