Joakimpalm-Zen/Qwen3-30B-A3B-selective-attnQ8_0-expQ4_0-GGUF
Text Generation • 31B • Updated • 419
GGUF models and adapters meant to be downloaded and served with Xyntetik Runner. Each card states its measured status against the parent.
Note Selective precision (attention Q8_0, experts Q4_0), 17.99 GB. Passes the fidelity bar where the uniform Q4_K_M fails it. The lead artifact.
Note Expert-pruned (120 of 128 kept), Q4_K_M, 17.5 GB. Passes both the original and the current bar.
Note Plain canonical Q8_0 of NVIDIA's base, the first Mamba-2 hybrid Runner serves. 5/6 greedy-identical vs llama.cpp.
Note The 2.63 GB quickstart file for an 8 GB machine. Fails the fidelity bar and the card says so.
Note LoRA trained through the served Q4_K_M weights, byte-reproducible. Held-out exact tool-call 0.69 base, 1.00 adapted.