遇见数据集

Medical-Reasoning-SFT-Qwen3-Next-80B

收藏
魔搭社区2026-07-13 更新2026-07-15 收录
官方服务:

资源简介:

# Medical-Reasoning-SFT-Qwen3-Next-80B A large-scale medical reasoning dataset generated using [Qwen/Qwen3-Next-80B-A3B-Thinking](https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Thinking), containing over 604,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. ## Dataset Overview | Metric | Value | |--------|-------| | **Model** | Qwen/Qwen3-Next-80B-A3B-Thinking | | **Total Samples** | 604,249 | | **Samples with Reasoning** | 604,249 (100%) | | **Estimated Tokens** | ~1.42 Billion | | **Content Tokens** | ~505 Million | | **Reasoning Tokens** | ~917 Million | | **Language** | English | ## Schema Each sample follows the conversational messages format with reasoning content: ```json { "messages": [ { "role": "system", "content": "You are a medical expert...", "reasoning_content": null }, { "role": "user", "content": "What are the symptoms of diabetes?", "reasoning_content": null }, { "role": "assistant", "content": "The main symptoms of diabetes include...", "reasoning_content": "Let me think through this systematically. Diabetes affects blood sugar regulation, so I should consider symptoms related to hyperglycemia..." } ] } ``` ### Fields | Field | Type | Description | |-------|------|-------------| | `messages` | list | Array of message objects in the conversation | | `messages[].role` | string | Either "system", "user", or "assistant" | | `messages[].content` | string | The main message content | | `messages[].reasoning_content` | string or null | Chain-of-thought reasoning (assistant messages only) | ## Usage ### Loading with Datasets Library ```python from datasets import load_dataset dataset = load_dataset("OpenMed/Medical-Reasoning-SFT-Qwen3-Next-80B") ``` ### Accessing Samples ```python # Get a sample sample = dataset['train'][0] # Access messages for msg in sample['messages']: print(f"Role: {msg['role']}") print(f"Content: {msg['content'][:100]}...") if msg['reasoning_content']: print(f"Reasoning: {msg['reasoning_content'][:100]}...") ``` ### Filtering by Reasoning ```python # Get samples with reasoning content samples_with_reasoning = dataset['train'].filter( lambda x: x['messages'][-1]['reasoning_content'] is not None ) ``` ## Model Characteristics Qwen3-Next-80B-A3B-Thinking is known for: - **Advanced Reasoning**: State-of-the-art thinking capabilities with dedicated reasoning architecture - **Large Scale**: 80B parameter model with extensive medical knowledge - **Detailed Traces**: Comprehensive step-by-step reasoning for complex medical problems - **High Quality**: Rigorous logical flow in medical explanations ## Intended Use This dataset is designed for: - **Fine-tuning medical reasoning models**: Train LLMs to provide detailed, step-by-step medical reasoning - **Chain-of-thought training**: Develop models that show their thinking process - **Medical QA systems**: Build question-answering systems for healthcare applications - **Research**: Study reasoning patterns in medical domain AI ## Limitations and Considerations - This dataset is generated by an AI model and should not be used as a substitute for professional medical advice - Responses may contain inaccuracies and should be validated by medical professionals - Not intended for clinical decision-making without expert review - The reasoning traces reflect the model's approach, not necessarily optimal clinical reasoning ## Citation If you use this dataset, please cite: ```bibtex @dataset{medical_reasoning_sft_qwen3_next_80b, title={Medical-Reasoning-SFT-Qwen3-Next-80B}, author={OpenMed}, year={2025}, publisher={Hugging Face}, url={https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Qwen3-Next-80B} } ``` ## License Apache 2.0

提供机构:
maas
创建时间:
2026-02-03
二维码
社区交流群
二维码
科研交流群
商业服务