遇见数据集

gemini-3.1-pro-hard-high-reasoning

收藏
魔搭社区2026-06-03 更新2026-07-15 收录
官方服务:

资源简介:

# Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M ## Dataset Details ### Dataset Description This dataset represents the frontier of synthetic reasoning data, generated by **[Gemini 3.1 Pro](https://deepmind.google/technologies/gemini/)** (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes **logical density** and **multi-step verification**. The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models, Gemini 3.1 Pro utilizes enhanced internal reasoning traces to tackle the most complex domains, from Quantum Field Theory to Zero-Knowledge Proofs, with far fewer hallucinations and significantly more robust derivations. - **Curated by:** Agentic workflow via Gemini 3 Flash & [Gemini 3.1 Pro](https://deepmind.google/technologies/gemini/) - **Total Token Volume:** 5.6 Million tokens - **Generation Cost:** $70.56 ($12.60 per 1M tokens) - **Primary Model:** [Gemini 3.1 Pro](https://deepmind.google/technologies/gemini/) (High Reasoning) - **License:** MIT ### Why Gemini 3.1 Pro (High Reasoning)? Compared to the 3.0 Pro model, the 3.1 Pro High Reasoning model excels in: 1. **Logical Persistence:** It maintains coherence over much longer derivation chains in STEM domains. 2. **Constraint Satisfaction:** Superior ability to follow "No Fluff" and "Niche Concept" requirements without defaulting to generic Wikipedia-style summaries. 3. **Recursive Self-Correction:** The 3.1 architecture is better at identifying and fixing its own logical errors in real-time before outputting the final answer. ## Uses ### Direct Use - **Frontier Model Training:** Ideal for fine-tuning 7B-70B models to reach "Expert-level" performance in technical sub-fields. - **Complexity Benchmarking:** Serving as a "Gold Standard" test set for other reasoning models (e.g., O1, Claude 3.5 Opus). - **Synthetic Reasoning Distillation:** Distilling the "High Reasoning" capabilities of 3.1 Pro into smaller, specialized agents. ## Dataset Creation ### Curation Rationale The dataset was designed to eliminate "retrieval-heavy" questions. If a question can be answered by a search engine, it was discarded. Every entry requires the model to synthesize information across at least three distinct logical steps. ### Source Data #### Data Collection and Processing The orchestrator (Gemini 3 Flash) used the following high-intensity system instruction to challenge the **[Gemini 3.1 Pro](https://deepmind.google/technologies/gemini/)** model: > **System Instruction:** > Act as a "Super-Intelligence Evaluator". You are generating training data to test the absolute limits of Gemini 3.1 Pro (current best high-reasoning model). > > **Task:** Generate {BATCH_SIZE} distinct, complex but solvable prompts/questions that require extreme logic. > **Target Domain:** {selected_domain} > > **Requirements:** > 1. **Difficulty:** The question must be unsolvable by simple retrieval. It requires multi-step logic, derivation, or synthesis of conflicting information. > 2. **Concept:** Pick a specific, niche concept within {selected_domain}. > 3. **Prompt Text:** The user prompt should be detailed (code snippets, math proofs, or philosophical paradoxes). > 4. **No Fluff:** Go straight to the hard part. > > Output strictly in JSON format. #### Domain Coverage (Expert-Level) The dataset spans 60+ "Hard" domains including: * **Physics:** QFT, General Relativity, Condensed Matter, Thermodynamics. * **Math:** Algebraic Topology, Analytic Number Theory, Category Theory, IMO Q3/Q6. * **Coding/CS:** Linux Kernel Dev, Rust Lifetimes, ZK-Proofs (SNARKs), LLVM IR. * **Biology/Med:** CRISPR Off-target Analysis, Protein Folding, T-Cell Signaling. * **Strategic Logic:** GTO Poker Solvers, Minimax AlphaBeta, Cryptanalysis. * **High-End Benchmarks:** ARC-AGI, GPQA Diamond, MMLU-Pro, HLE. ## Dataset Structure | Field | Type | Description | |---|---|---| | `domain` | String | The specialized field (e.g., "Analytic Number Theory") | | `difficulty` | String | Fixed to: Extreme, World-Class, PhD-Level, or Grandmaster | | `reasoning_complexity` | Integer | A score of 1-10 on the depth of the logic chain | | `prompt` | String | The ultra-complex technical query | | `high_reasoning_output` | String | The verified solution provided by Gemini 3.1 Pro | ## Bias, Risks, and Limitations - **Expertise Requirement:** The solutions are so technical that they require subject-matter experts (PhDs) to verify accuracy. - **Compute Intensity:** Training on this data requires models with high context windows to accommodate the long reasoning traces. ## Citation **BibTeX:** ```bibtex @dataset{gemini_3_1_pro_high_reasoning_5_6m, author = {Gemini 3 Flash & Gemini 3.1 Pro}, title = {Gemini 3.1 Pro Ultra-Reasoning Dataset (5.6M Tokens)}, year = {2026}, publisher = {HuggingFace}, url = {https://deepmind.google/technologies/gemini/} }

提供机构:
maas
创建时间:
2026-04-07
二维码
社区交流群
二维码
科研交流群
商业服务