SINAI/ALIA-es-cultural-heritage-synthetic-instructions
收藏资源简介:
ALIA西班牙文化与遗产合成指令语料库是一个西班牙语合成指令调优资源,基于ALIA项目采用Magpie方法创建。该数据集旨在训练和评估语言模型在文化遗产、数字人文和历史知识任务中的表现,具有自然语言变体和大规模监督。它包含748,480个实例,629,682,398个令牌,涵盖25种任务模态,包括遗产问答、真/假判断、主题/流派/时期分类、实体提取、摘要、简化以及档案分析等。数据生成过程遵循Magpie方法,模型首先生成用户侧查询(遗产问题、历史探究或档案任务),然后生成相应答案。语料库包含基于西班牙文化遗产文档(历史文本、博物馆目录、档案描述、学术遗产出版物)构建的规范性和自然变体提示,涵盖多样任务类型,如遗产问答、文化分类、实体提取、文本生成以及历史和遗产内容分析,并包含多种语言风格变体(正式、口语化、感叹式、电报式)。
The ALIA Spanish Cultural and Heritage Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in cultural heritage, digital humanities, and historical knowledge tasks with natural linguistic variation and large-scale supervision. It contains 748,480 instances, 629,682,398 tokens, and 25 task modalities (heritage QA, true/false, topic/genre/period classification, entity extraction, summarization, simplification, and archival analysis). The dataset contains synthetic cultural heritage and history instruction-response pairs generated with instruction models and curated through a multi-step quality pipeline. The generation process follows Magpie, where the model first produces a user-side query (heritage question, historical inquiry, or archival task) and then generates the corresponding answer. The corpus includes both canonical and naturally varied prompts built from Spanish cultural heritage documents (historical texts, museum catalogs, archival descriptions, academic heritage publications). It includes diverse task types such as heritage QA, cultural classification, entity extraction, text generation, and analysis of historical and patrimonial content, with multiple linguistic style variations (formal, colloquial, exclamatory, telegraphic).




