遇见数据集

gemini-3-pro-10000x-hard-high-reasoning

收藏
魔搭社区2026-05-23 更新2026-07-15 收录
官方服务:

资源简介:

# Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning ## Dataset Details ### Dataset Description ### Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement. This dataset is a high-complexity synthetic reasoning corpus containing expert-level problems and solutions. It was generated to test the absolute limits of modern reasoning models, focusing on domains requiring multi-step logic, derivation, and synthesis of conflicting information. The dataset was constructed using an agentic workflow where **Gemini 3 Flash** acted as the prompt engineer/orchestrator to generate distinct, complex scenarios, which were then solved by **[Gemini 3 Pro](https://deepmind.google/technologies/gemini/pro/)**. - **Curated by:** Synthetic generation via Gemini 3 Flash & Gemini 3 Pro - **Total Token Volume:** 17.8 Million tokens - **Creation Cost:** ~$224.23 USD - **Language(s):** English (Technical/Academic) - **License:** MIT but please work with me if you want to use the dataset. ### Dataset Sources - **Generator Model:** [Gemini 3 Pro](https://deepmind.google/technologies/gemini/pro/) (Reasoning/Response) - **Orchestrator Model:** Gemini 3 Flash and GPT 5.2 (Prompt Generation) ## Uses ### Direct Use - **SFT (Supervised Fine-Tuning):** Enhancing the reasoning capabilities of smaller models in niche domains (Quantum Physics, Kernel Development, Algebraic Topology). - **Evaluation:** Benchmarking models against "PhD-Level" and "Grandmaster" difficulty tasks. - **RAG Testing:** Evaluating retrieval systems on "unsolvable by simple retrieval" queries. ### Out-of-Scope Use - **General Chat:** The data is highly technical and dense; not suitable for casual conversation tuning. - **Basic Instruction Following:** This dataset focuses on complex reasoning chains, not simple formatting or summarization tasks. ## Dataset Structure The dataset follows a structured format containing the domain metadata, the difficulty rating, the complex prompt, and the reasoning trace/solution. | Field | Description | |---|---| | `domain` | The specific academic or technical field (e.g., "Quantum Field Theory", "Rust Memory Safety"). | | `difficulty` | Categorized as "Extreme", "World-Class", "PhD-Level", or "Grandmaster". | | `topic` | The specific sub-concept (e.g., "Riemann Zeta Function", "Paxus Consensus"). | | `prompt` | The complex scenario generated by Gemini 3 Flash. | | `response` | The solution or derivation provided by Gemini 3 Pro. | ## Dataset Creation ### Curation Rationale The goal was to move beyond standard benchmarks by synthesizing data that targets the specific "hard parts" of various disciplines. The curation logic enforces that questions **cannot be solved by simple retrieval** and must require logic or calculation. ### Source Data #### Data Collection and Processing The data was generated using a high-throughput pipeline: 1. **Prompt Engineering (Gemini 3 Flash):** Gemini 3 Flash was instructed to act as a "Super-Intelligence Evaluator" to generate prompts. *System Instruction used:* > Act as a "Super-Intelligence Evaluator". You are generating training data to test the absolute limits of Gemini 3 pro (current best reasoning model). > > **Task:** Generate {BATCH_SIZE} distinct, complex but solvable prompts/questions that requires really strong logic. > **Target Domain:** {selected_domain} > > **Requirements:** > 1. **Difficulty:** The question must be unsolvable by simple retrieval. It requires multi-step logic, derivation, or synthesis of conflicting information. > 2. **Concept:** Pick a specific, niche concept within {selected_domain}. (e.g., if Physics, don't ask about gravity; ask about the stress-energy tensor in a specific metric). > 3. **Prompt Text:** The user prompt (text) should be detailed. It can be a scenario, a code snippet to debug, a math problem, or a philosophical paradox. > 4. **No Fluff:** Go straight to the hard part. > > If the domain is code, provide a complex prompt to implement or a broken piece of complex code or a system design requirement that is notoriously hard. > If the domain is math, ensure it is proof-based or calculation-heavy (AIME/Putnam level). > > Output strictly in JSON format matching the schema. 2. **Response Generation ([Gemini 3 Pro](https://deepmind.google/technologies/gemini/pro/)):** The resulting prompts were fed to Gemini 3 Pro to generate high-fidelity solutions. #### Domain Coverage The dataset covers a massive range of disciplines, specifically targeting hard benchmarks: * **LLM Benchmarks:** GPQA Diamond, MMLU-Pro, BBH, ARC-AGI, MathVista. * **Systems Programming:** Linux Kernel (C/Assembly), Rust Memory Safety, Distributed Systems (Paxos/Raft), CUDA, Reverse Engineering. * **Advanced Mathematics:** AIME 2026, Putnam, Algebraic Topology, Stochastic Calculus, Differential Geometry. * **Physics & Chemistry:** QFT, General Relativity, Condensed Matter, Organic Retro-Synthesis. * **Biology & Medicine:** USMLE Step 3 (Complex Cases), CRISPR Analysis, AlphaFold/Protein Folding. * **Law & Humanities:** International Arbitration, Patent Law, Continental Philosophy, Heterodox Economics. * **Logic & Strategy:** Chess Engines (Minimax), Poker GTO, Cryptanalysis, CTF Web Security. ### Cost & Compute - ### Total Tokens Generated:**17.8 Million** - ### Estimated Cost:** 12.6$/m * 17.8m = **$ 224.28** - ### Efficiency:** Leveraged Gemini 3 Flash for efficient prompt scaffolding and [Gemini 3 Pro](https://deepmind.google/technologies/gemini/pro/) for reasoning density. ## Bias, Risks, and Limitations - **Jialbreak** While Gemini 3 Pro is highly capable, it may jailbreak the model - but can lead in improvement, as now model is unrestricted, some hard, logical "not-safe" questions were FULLY answered with system instructions that mention to answer for educational purpose. - **Code Safety:** The dataset contains examples of "Reverse Engineering" and "Exploit generation" (CTF/Security) for educational and defensive purposes. Users should execute generated code in sandboxed environments. - **Difficulty:** This dataset is intentionally skewed towards "PhD-Level" problems; but i added general purpose samples - so it is suitable for fine tuning general purpose models

提供机构:
maas
创建时间:
2026-04-07
二维码
社区交流群
二维码
科研交流群
商业服务