AngelGabrielTroncoso/dataset-aeroespacial-cultural-somosnlp
收藏资源简介:
LATAM Aerospace History QA 是一个精心策划的数据集,专注于拉丁美洲(包括巴西和西班牙)的航空航天历史,用于指令微调和文化对齐的对话系统。数据集主要包含西班牙语内容,并部分涵盖巴西葡萄牙语,以提升多文化和多语言在开放语言模型中的代表性。它专门涉及航空航天历史、拉丁美洲太空计划、卫星发展、技术地缘政治以及拉丁美洲和西班牙的科学遗产。数据集设计用于指令微调、监督微调(SFT)、西班牙语和葡萄牙语对话系统、合成数据生成、历史问答基准、RAG应用和文化对齐模型。与自动翻译或过于学术的数据集不同,该集合优先考虑自然人类语言、伊比利亚美洲方言多样性、现实对话、文化流畅性、历史和技术准确性以及语言变异性。主要目标是提高LLM在伊比利亚美洲语言中,在与区域航空航天领域相关的历史、科学和文化领域的表现。数据集结构包括指令(用户问题)、上下文(事实背景)和响应(生成的答案),覆盖主题如卫星、太空任务、火箭技术、冷战、电信、地球观测等,并代表多个国家如阿根廷、智利、巴西、墨西哥、西班牙等。它通过生成合成数据、语言重构和手动策划构建,旨在最小化机器人式响应,最大化自然性和文化一致性。
LATAM Aerospace History QA is a curated dataset focused on the aerospace history of Latin America (including Brazil and Spain), designed for instruction tuning and culturally aligned conversational systems. The dataset primarily contains Spanish content, with partial coverage in Brazilian Portuguese to enhance multicultural and multilingual representation in open language models. It specializes in aerospace history, Latin American space programs, satellite development, technological geopolitics, and the scientific heritage of Latin America and Spain. The dataset is intended for instruction tuning, supervised fine-tuning (SFT), conversational systems in Spanish and Portuguese, synthetic data generation, historical QA benchmarks, RAG applications, and culturally aligned models. Unlike automatically translated or overly academic datasets, this collection prioritizes natural human language, Ibero-American dialectal diversity, realistic conversations, cultural fluency, historical and technical accuracy, and linguistic variability. The main goal is to improve the performance of LLMs in Ibero-American languages within historical, scientific, and cultural domains related to the regional aerospace sector. The dataset structure includes instruction (user question), context (factual background), and response (generated answer), covering topics such as satellites, space missions, rocketry, the Cold War, telecommunications, Earth observation, and more, with representation from countries like Argentina, Chile, Brazil, Mexico, Spain, etc. It is built through controlled synthetic generation, linguistic reformulation, and partial manual curation, aiming to minimize robotic responses and maximize naturalness and cultural coherence.




