Instructions to use Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF:Q8_0
Use Docker
docker model run hf.co/Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF:Q8_0
- LM Studio
- Jan
- vLLM
How to use Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF:Q8_0
- Ollama
How to use Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF with Ollama:
ollama run hf.co/Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF:Q8_0
- Unsloth Desktop
- Pi
How to use Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF with Docker Model Runner:
docker model run hf.co/Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF:Q8_0
- Lemonade
How to use Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF:Q8_0
Run and chat with the model
lemonade run user.NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
NVIDIA Nemotron-Nano-9B-v2 — Q8_0 GGUF (canonical, xyntetik-runner)
Release: built to be downloaded and served with Xyntetik Runner.
- Measured status against the parent: against NVIDIA's own float32 implementation, 100% margin-qualified top-1 at Q8_0 over 100 positions (llama.cpp also 100%; Runner's mean KL the lower, 0.0041 against 0.0054). Second column: a plain near-lossless Q8_0 by Runner's own canonical quantiser, 5 of 6 greedy outputs identical to llama.cpp b10353 (the miss a near-tie). Not a bar-gated derivative.
- Parent model: nvidia/NVIDIA-Nemotron-Nano-9B-v2, NVIDIA Open Model License
- Collection: Runner Releases, also in Runner Compatibility Reports
- Evidence: the publisher-reference table on this card (golden pass, 2026-09-07), then the greedy-identity checks
- Runner compatibility: the first Mamba-2 hybrid (
nemotron_h) Runner serves; CPU path for the recurrent layers.
Want this for another model? The tools behind these files are open: requantization, expert pruning and sublayer removal in Xyntetik Runner, and the method for this file is on this card or in its evidence dataset.
Request a model · Report your result · Community results · All files
If this file is useful to you, a contribution funds the hardware time behind the next one.
Serve with xyntetik-runner — tool calls that still parse when the budget runs out
This GGUF is built to be served with xyntetik-runner, a single-binary, dependency-free C11 inference engine (CPU / CUDA / Metal, OpenAI-compatible server). Its one differentiator you cannot get from another engine:
When a tool call runs past its token budget, Runner closes it to the smallest schema-legal document, so the arguments still parse.
Most OpenAI-compatible engines return finish_reason: "length" with an empty or malformed tool_calls the caller has to repair or retry. Runner does forced-truncation recovery: once a JSON document starts, it emits a legal ending when the budget expires. On local models — tight context, slow generation — that is the difference between an agent loop that finishes and one that retries from scratch.
Measured — what each engine hands the caller when the token budget cuts a tool call short (same box, same tool schema, same prompt, tool_choice:"required", temperature 0, budgets 1→64):
| engine | budget too small (1–16 tokens) | enough budget (64, control) |
|---|---|---|
| Runner | executable tool_calls, arguments parse |
completes |
| vLLM 0.27.1 | no call; protocol framing leaks into content |
completes |
| llama.cpp b10488 | no call; leak, then tool_calls with unparseable args |
completes |
| Ollama 0.32.14 | no call; empty content, then HTTP 500 | completes |
| TensorRT-LLM 1.2.1 †| no call; <tool_call> leak, then empty content |
completes |
| SGLang 0.5.17 †| no call; <tool_call> leak, then empty content |
completes |
Every engine completes at 64 tokens — the failure is truncation, not misconfiguration. Below that, only Runner returns an executable call. This is the behaviour across every OpenAI-compatible engine measured, not a claim about engines not measured. †TensorRT-LLM and SGLang were measured on a Qwen3-1.7B substitute; truncation recovery is a property of the runtime, not the model.
Constrained decoding also holds the shape of a tool call under quantization: on a full quant ladder, schema conformance and tool selection stayed 100% down to Q4_0 (argument agreement decayed to 50%) — it guarantees the shape of a call at any quantization, not its contents.
Quickstart
# build the single binary (no deps), then serve this GGUF:
runner -m NVIDIA-Nemotron-Nano-9B-v2-Q8_0.gguf --serve
# point any OpenAI-compatible client at http://localhost:8080/v1
Full recipe, raw responses, and the per-release regression gate: the truncation benchmark.
Runs on xyntetik-runner (nemotron_h hybrid)
Nemotron-Nano-9B-v2 is a hybrid Mamba-2 / attention model — architecture nemotron_h, a state-space (SSM) + attention hybrid. This GGUF loads and decodes on xyntetik-runner (CPU / CUDA) and on stock llama.cpp (verified against b10353).
runner -m NVIDIA-Nemotron-Nano-9B-v2-Q8_0.gguf --serve
Canonical, interoperable Q8_0
This file was produced by xyntetik-runner --quantize --quant q8_0 from the bf16 GGUF base. After a quantizer fix (efb041a), runner's Q8_0 tensor data is byte-for-byte identical to what llama.cpp's own Q8_0 quantizer produces from the same input — all 122 quantized tensors and all 219 F32 tensors match exactly. Runner emits canonical ggml Q8_0: a runner-produced Q8_0 and a llama.cpp-produced Q8_0 of the same input are the same weights (only container metadata KV ordering differs).
Fidelity vs the bf16 base — measured, not asserted
Against the publisher's reference first. The golden pass of 2026-09-07 (report) scored both engines against NVIDIA's own implementation (transformers 5.16, NemotronHForCausalLM, float32, eager attention, CPU) on 100 corpus positions, with llama.cpp 73a43d1 served beside Runner on the same files. Margin-qualified top-1 is the share of positions where the reference had a clear preference (top two more than 0.5 nat apart) and the engine picked it; it is the measure this project gates on, so it comes first.
| weights | engine | margin-qualified top-1 | top-1 | mean KL from the reference |
|---|---|---|---|---|
| bf16 | Runner | 100% | 98% | 0.0027 |
| bf16 | llama.cpp | 100% | 98% | 0.0031 |
| Q8_0 (this file name) | Runner | 100% | 97% | 0.0041 |
| Q8_0 (this file name) | llama.cpp | 100% | 95% | 0.0054 |
At 100 positions the two engines are indistinguishable on the argmax; the KL difference is small and is the measure that favours Runner, so it is listed last. The 2,000-position rerun that resolved the other families did not finish for this one. Everything below compares against the bf16 base or against llama.cpp, which is agreement with another engine, not fidelity to the publisher.
Two independent CPU/greedy checks against the bf16 base as ground truth:
Next-token fidelity — 400-position KLD + top-1 agreement over a pinned corpus (top-20 logprob surface, scripts/kld-compare.py):
| Q8_0 file | mean KLD vs bf16 | top-1 agreement | top-8 overlap |
|---|---|---|---|
| this file (canonical runner Q8_0) | 2.46e-4 | 100.0% | 97.7% |
| llama.cpp Q8_0 (bartowski) | 2.46e-4 | 100.0% | 97.7% |
| pre-fix runner Q8_0 (non-canonical) | 2.41e-4 | 100.0% | 97.8% |
All three are within noise of each other and agree with bf16 on the top-1 token 100% of the time — Q8_0 is effectively lossless here. The canonical file's KLD is identical to llama.cpp's (byte-identical weights). The pre-fix, non-canonical runner file was not lower quality — it was merely bit-different from ggml; the fix is an interoperability / canonicality win, not a quality change.
Server-path greedy identity (2026-08-20, reproducible) — via the committed scripts/reference_compare.py (both engines served, exact generated UTF-8 from /v1/completions at temperature 0; 5 prompts × 32 tokens, CPU): 5/5 byte-identical vs llama.cpp ea12b27 — factual, code, counting, story, and JSON prompts all match exactly. Raw JSON: ssm-greedy-reference-2026-08-20/.
Byte-identity spot-check — 6-prompt greedy vs the bf16 base, on both engines: this Q8_0 scores 5/6 on runner and 5/6 on llama.cpp. The one miss is a near-tie argmax coin-flip (identical prefix, one differing word choice) — and the canonical llama.cpp Q8_0 shows the same miss. This is a near-tie byte-identity check, not a quality measurement.
An envelope manifest (NVIDIA-Nemotron-Nano-9B-v2-Q8_0.gguf.envelope.json) is included. Verdict: experimental — a dated measurement, not a standing quality claim (no per-model conformance gate report).
Model, attribution, and license
- Base model: nvidia/NVIDIA-Nemotron-Nano-9B-v2, a hybrid Mamba-2 / attention (
nemotron_h) model by NVIDIA. All model weights, capabilities, and credit belong to NVIDIA; this repository only redistributes a quantized GGUF conversion of NVIDIA's weights. - License: NVIDIA Open Model License. Use of this model is governed by the NVIDIA Open Model License; by using these weights you agree to its terms.
- Quantization: bf16 GGUF → Q8_0 via
xyntetik-runner --quantize --quant q8_0(canonical ggml Q8_0). GGUF sha256d3224286484df1ae79c2b74d2a0913e1bbb5a364b0dd7c1061c95fa97ee608d7, ~8.81 GiB (9,458,093,632 bytes). - Not affiliated with, sponsored by, or endorsed by NVIDIA.
- Downloads last month
- 354
8-bit
Model tree for Joakimpalm-Zen/NVIDIA-Nemotron-Nano-9B-v2-Q8_0-GGUF
Base model
nvidia/NVIDIA-Nemotron-Nano-12B-v2-Base