WithinUsAI/The_Research_From_WithIn_10k
收藏资源简介:
The_Research_From_WithIn_10k 是一个前沿质量的数据集,旨在训练高级自主代理语言模型掌握高质量的研究方法、信息收集、来源评估、证据合成、知识整合、批判性分析和专业报告。每个示例都提供一个真实的专业研究任务,包括明确的目标、聚焦的研究问题、从多个来源收集的证据、来源评估、综合发现、识别的知识差距和基于证据的结论。思考轨迹展示了专业调查、证据收集、合成、规划、优先级排序、评估、验证和决策制定。数据集包含10,000个独特示例,覆盖AI研究、机器学习、软件工程、云计算、网络安全防御、医学、生物学、化学、物理学、经济学、金融、初创企业、业务运营、基础设施、公共政策、教育、制造业、物流、机器人技术和科学发现等多个领域。数据格式为JSONL,每个JSON对象包含指令、输入和输出字段,输出字段进一步细分为思考、研究目标、研究问题、收集的证据、来源评估、综合发现、知识差距和结论。数据集经过严格的去重和验证过程,确保逻辑一致性和专业质量,适用于监督微调代理模型进行研究和证据合成任务。
The_Research_From_WithIn_10k is a cutting-edge high-quality dataset designed to train advanced autonomous agent large language models to master high-quality research methodologies, information gathering, source evaluation, evidence synthesis, knowledge integration, critical analysis, and professional report writing. Each sample provides a real-world professional research task, including clear objectives, focused research questions, evidence collected from multiple sources, source evaluation, synthesized findings, identified knowledge gaps, and evidence-based conclusions. Thought trajectories demonstrate professional investigation, evidence collection, synthesis, planning, prioritization, evaluation, validation, and decision-making. The dataset contains 10,000 unique samples covering multiple domains including AI research, machine learning, software engineering, cloud computing, cybersecurity defense, medicine, biology, chemistry, physics, economics, finance, startups, business operations, infrastructure, public policy, education, manufacturing, logistics, robotics, and scientific discovery. The data format is JSONL, with each JSON object containing fields for "instruction", "input", and "output"; the "output" field is further subdivided into "thought", "research objective", "research question", "collected evidence", "source evaluation", "synthesized findings", "knowledge gaps", and "conclusion". The dataset has undergone strict deduplication and validation processes to ensure logical consistency and professional quality, making it suitable for supervised fine-tuning of agent models for research and evidence synthesis tasks.




