NVIDIA Nemotron-Nano-9B-v2 — Q8_0 GGUF (canonical, xyntetik-runner)

Release: built to be downloaded and served with Xyntetik Runner.

  • Measured status against the parent: against NVIDIA's own float32 implementation, 100% margin-qualified top-1 at Q8_0 over 100 positions (llama.cpp also 100%; Runner's mean KL the lower, 0.0041 against 0.0054). Second column: a plain near-lossless Q8_0 by Runner's own canonical quantiser, 5 of 6 greedy outputs identical to llama.cpp b10353 (the miss a near-tie). Not a bar-gated derivative.
  • Parent model: nvidia/NVIDIA-Nemotron-Nano-9B-v2, NVIDIA Open Model License
  • Collection: Runner Releases, also in Runner Compatibility Reports
  • Evidence: the publisher-reference table on this card (golden pass, 2026-09-07), then the greedy-identity checks
  • Runner compatibility: the first Mamba-2 hybrid (nemotron_h) Runner serves; CPU path for the recurrent layers.

Want this for another model? The tools behind these files are open: requantization, expert pruning and sublayer removal in Xyntetik Runner, and the method for this file is on this card or in its evidence dataset.

Request a model · Report your result · Community results · All files

Support this work

If this file is useful to you, a contribution funds the hardware time behind the next one.

Serve with xyntetik-runner — tool calls that still parse when the budget runs out

This GGUF is built to be served with xyntetik-runner, a single-binary, dependency-free C11 inference engine (CPU / CUDA / Metal, OpenAI-compatible server). Its one differentiator you cannot get from another engine:

When a tool call runs past its token budget, Runner closes it to the smallest schema-legal document, so the arguments still parse.

Most OpenAI-compatible engines return finish_reason: "length" with an empty or malformed tool_calls the caller has to repair or retry. Runner does forced-truncation recovery: once a JSON document starts, it emits a legal ending when the budget expires. On local models — tight context, slow generation — that is the difference between an agent loop that finishes and one that retries from scratch.

Measured — what each engine hands the caller when the token budget cuts a tool call short (same box, same tool schema, same prompt, tool_choice:"required", temperature 0, budgets 1→64):

engine budget too small (1–16 tokens) enough budget (64, control)
Runner executable tool_calls, arguments parse completes
vLLM 0.27.1 no call; protocol framing leaks into content completes
llama.cpp b10488 no call; leak, then tool_calls with unparseable args completes
Ollama 0.32.14 no call; empty content, then HTTP 500 completes
TensorRT-LLM 1.2.1 † no call; <tool_call> leak, then empty content completes
SGLang 0.5.17 † no call; <tool_call> leak, then empty content completes

Every engine completes at 64 tokens — the failure is truncation, not misconfiguration. Below that, only Runner returns an executable call. This is the behaviour across every OpenAI-compatible engine measured, not a claim about engines not measured. † TensorRT-LLM and SGLang were measured on a Qwen3-1.7B substitute; truncation recovery is a property of the runtime, not the model.

Constrained decoding also holds the shape of a tool call under quantization: on a full quant ladder, schema conformance and tool selection stayed 100% down to Q4_0 (argument agreement decayed to 50%) — it guarantees the shape of a call at any quantization, not its contents.

Quickstart

# build the single binary (no deps), then serve this GGUF:
runner -m NVIDIA-Nemotron-Nano-9B-v2-Q8_0.gguf --serve
# point any OpenAI-compatible client at http://localhost:8080/v1

Full recipe, raw responses, and the per-release regression gate: the truncation benchmark.

Runs on xyntetik-runner (nemotron_h hybrid)

Nemotron-Nano-9B-v2 is a hybrid Mamba-2 / attention model — architecture nemotron_h, a state-space (SSM) + attention hybrid. This GGUF loads and decodes on xyntetik-runner (CPU / CUDA) and on stock llama.cpp (verified against b10353).

runner -m NVIDIA-Nemotron-Nano-9B-v2-Q8_0.gguf --serve

Canonical, interoperable Q8_0

This file was produced by xyntetik-runner --quantize --quant q8_0 from the bf16 GGUF base. After a quantizer fix (efb041a), runner's Q8_0 tensor data is byte-for-byte identical to what llama.cpp's own Q8_0 quantizer produces from the same input — all 122 quantized tensors and all 219 F32 tensors match exactly. Runner emits canonical ggml Q8_0: a runner-produced Q8_0 and a llama.cpp-produced Q8_0 of the same input are the same weights (only container metadata KV ordering differs).

Fidelity vs the bf16 base — measured, not asserted

Against the publisher's reference first. The golden pass of 2026-09-07 (report) scored both engines against NVIDIA's own implementation (transformers 5.16, NemotronHForCausalLM, float32, eager attention, CPU) on 100 corpus positions, with llama.cpp 73a43d1 served beside Runner on the same files. Margin-qualified top-1 is the share of positions where the reference had a clear preference (top two more than 0.5 nat apart) and the engine picked it; it is the measure this project gates on, so it comes first.

weights engine margin-qualified top-1 top-1 mean KL from the reference
bf16 Runner 100% 98% 0.0027
bf16 llama.cpp 100% 98% 0.0031
Q8_0 (this file name) Runner 100% 97% 0.0041
Q8_0 (this file name) llama.cpp 100% 95% 0.0054

At 100 positions the two engines are indistinguishable on the argmax; the KL difference is small and is the measure that favours Runner, so it is listed last. The 2,000-position rerun that resolved the other families did not finish for this one. Everything below compares against the bf16 base or against llama.cpp, which is agreement with another engine, not fidelity to the publisher.

Two independent CPU/greedy checks against the bf16 base as ground truth:

Next-token fidelity — 400-position KLD + top-1 agreement over a pinned corpus (top-20 logprob surface, scripts/kld-compare.py):

Q8_0 file mean KLD vs bf16 top-1 agreement top-8 overlap
this file (canonical runner Q8_0) 2.46e-4 100.0% 97.7%
llama.cpp Q8_0 (bartowski) 2.46e-4 100.0% 97.7%
pre-fix runner Q8_0 (non-canonical) 2.41e-4 100.0% 97.8%

All three are within noise of each other and agree with bf16 on the top-1 token 100% of the time — Q8_0 is effectively lossless here. The canonical file's KLD is identical to llama.cpp's (byte-identical weights). The pre-fix, non-canonical runner file was not lower quality — it was merely bit-different from ggml; the fix is an interoperability / canonicality win, not a quality change.

Server-path greedy identity (2026-08-20, reproducible) — via the committed scripts/reference_compare.py (both engines served, exact generated UTF-8 from /v1/completions at temperature 0; 5 prompts × 32 tokens, CPU): 5/5 byte-identical vs llama.cpp ea12b27 — factual, code, counting, story, and JSON prompts all match exactly. Raw JSON: ssm-greedy-reference-2026-08-20/.

Byte-identity spot-check — 6-prompt greedy vs the bf16 base, on both engines: this Q8_0 scores 5/6 on runner and 5/6 on llama.cpp. The one miss is a near-tie argmax coin-flip (identical prefix, one differing word choice) — and the canonical llama.cpp Q8_0 shows the same miss. This is a near-tie byte-identity check, not a quality measurement.

An envelope manifest (NVIDIA-Nemotron-Nano-9B-v2-Q8_0.gguf.envelope.json) is included. Verdict: experimental — a dated measurement, not a standing quality claim (no per-model conformance gate report).

Model, attribution, and license

  • Base model: nvidia/NVIDIA-Nemotron-Nano-9B-v2, a hybrid Mamba-2 / attention (nemotron_h) model by NVIDIA. All model weights, capabilities, and credit belong to NVIDIA; this repository only redistributes a quantized GGUF conversion of NVIDIA's weights.
  • License: NVIDIA Open Model License. Use of this model is governed by the NVIDIA Open Model License; by using these weights you agree to its terms.
  • Quantization: bf16 GGUF → Q8_0 via xyntetik-runner --quantize --quant q8_0 (canonical ggml Q8_0). GGUF sha256 d3224286484df1ae79c2b74d2a0913e1bbb5a364b0dd7c1061c95fa97ee608d7, ~8.81 GiB (9,458,093,632 bytes).
  • Not affiliated with, sponsored by, or endorsed by NVIDIA.
Downloads last month
354
GGUF
Model size
9B params
Architecture
nemotron_h
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF

Collections including Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF