遇见数据集

ShanzaGull/world_Top_leaders_Dataset

收藏
Hugging Face2026-04-27 更新2026-05-03 收录
官方服务:

资源简介:

World Leaders ChatML数据集是一个合成生成的数据集,包含大约2000个问答对,专注于15位极具影响力的世界领导人的生活、成就和历史意义。该数据集专门设计用于使用标准的ChatML对话格式对大型语言模型(如TinyLlama、Llama 2或Mistral)进行指令微调和微调。支持的任务包括指令微调(训练模型响应历史和传记查询)和对话问答(构建作为历史专家或教育导师的AI助手)。数据集以JSONL格式提供,每个条目表示一个对话轮次,包含角色(用户或助手)和内容字段。数据来源于维基百科页面的自动爬取,经过清理、分块和生成过程,涵盖的领导者包括纳尔逊·曼德拉、温斯顿·丘吉尔等15位人物。当前版本包含占位符答案,需替换为高质量LLM生成的答案以用于实际训练。

The World Leaders ChatML Dataset is a synthetically generated dataset comprising approximately 2,000 question-and-answer pairs focused on the lives, accomplishments, and historical significance of 15 highly influential world leaders. This dataset is specifically designed for instruction-tuning and fine-tuning Large Language Models (LLMs), such as TinyLlama, Llama 2, or Mistral, using the standard ChatML conversational format. Supported tasks include Instruction Fine-Tuning (training models to respond to historical and biographical queries) and Conversational QA (building AI assistants that act as historical experts or educational tutors). The dataset is provided in a .jsonl format, with each entry representing a conversation turn containing role (user or assistant) and content fields. Data is sourced from automatically scraped Wikipedia pages, processed through cleaning, chunking, and generation steps, covering leaders such as Nelson Mandela, Winston Churchill, and others. The current iteration contains placeholder answers that should be replaced with high-quality LLM-generated responses for production training.

提供机构:
ShanzaGull
二维码
社区交流群
二维码
科研交流群
商业服务