遇见数据集

nemotron-super-reap-artifacts-draft

收藏
魔搭社区2026-05-02 更新2026-07-15 收录
官方服务:

资源简介:

> [!TIP] > Support this work: **[donate.sybilsolutions.ai](https://donate.sybilsolutions.ai)** > > REAP surfaces: [GLM](https://huggingface.co/spaces/0xSero/reap-glm-family) | [MiniMax](https://huggingface.co/spaces/0xSero/reap-minimax-family) | [Qwen](https://huggingface.co/spaces/0xSero/reap-qwen-family) | [Gemma](https://huggingface.co/spaces/0xSero/reap-gemma-family) | [Paper](https://arxiv.org/abs/2510.13999) | [Code](https://github.com/CerebrasResearch/reap) | [PR17](https://github.com/CerebrasResearch/reap/pull/17) | [Cerebras Collection](https://huggingface.co/collections/cerebras/cerebras-reap) # Nemotron Super REAP artifacts draft This dataset repo is a **draft research artifact release** for REAP-based observation and compression work on [nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16). ## Provenance - Upstream base model: [nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16) - This repo does **not** republish the original NVIDIA base weights. - This repo contains observation outputs, rankings, heatmaps, residency planning artifacts, and compression metadata derived from that base model. - The derived checkpoint weights live in separate draft model repos. ## What this contains - long-lane and short-lane validation/manifests/summaries - merged observation outputs and validation - rankings, analysis, heatmaps, budget sweep, and residency plans - compression summaries and smoke-test artifacts for the 25 percent and 50 percent REAP-pruned variants ## Method summary - Model under study: `NVIDIA-Nemotron-3-Super-120B-A12B-BF16` - Architecture observed at runtime: `NemotronHForCausalLM` - Hybrid depth pattern: `88` total blocks = `40` Mamba + `40` MoE + `8` attention - Routed experts per MoE layer: `512` - Experts per token: `22` - Observation method: REAP **layerwise** MoE observer - Long lane: `nemotron_super_long50_16k_v3` - `50` trajectories - cap `16384` tokens - `819200` total tokens - Short lane: `nemotron_super_short_mix_15120_t1024_b8192_v4` - mixed personal + bounded public prompts - cap `1024` tokens - `321560` total tokens - Canonical merged lane: `nemotron_super_merged_long50_short15120_v2` - `1140760` total tokens - `40` observed MoE layers - `512` experts per layer - exact routed accounting: `22.0` experts per token ## Draft status and limits - This is a **draft release**. - The artifacts are intended for systems research, residency planning, and follow-on offload experiments. - No claim is made here that these artifacts alone establish end-user quality or production readiness. - Serving benchmarks and AutoRound quantization publication are separate workstreams. ## License and terms Use of the upstream model and any derivative weights remains governed by the NVIDIA Open Model License included in `LICENSE.txt`. See the upstream base model card for additional terms and disclosures.

提供机构:
maas
创建时间:
2026-03-20
二维码
社区交流群
二维码
科研交流群
商业服务