Qwen3.8-Flash-Next — Abliterated · HGN-HQ (λ=1.5)

Downloads

HGN-HQ quantized build of windowsxp811203/Qwen3.8-Flash-Next-Abliterated for halogen-flash-server and the gfx1151-engine runtime.

This is not a transformers / vLLM / llama.cpp checkpoint. There are no safetensors or GGUF files here; the .hgn set only loads in an HGN-compatible engine. For the bf16 weights, use the source repo above.

What this build is

The source checkpoint removes refusal by orthogonalizing 151 residual-stream-writing tensors against a captured direction at λ=1.5 (W_ablit = (I − λ r rᵀ) W_stock). Naively quantizing those edited tensors would blur exactly the edit the model exists for — so this build quantizes around the direction and then repairs the experts against it:

  1. Direction recovery. The shared refusal direction r is recovered from the official Halogen quality overlay (qwen38-flash-next-w4b.overlay.hgn) and the abliterated BF16 checkpoint, using 12 anchor tensors (all self_attn.o_proj / linear_attn.out_proj / shared_expert.down_proj spread across the 48 layers). Recovery was effectively exact: min anchor cosine 0.99995, median s1/s2 ≈ 57, per-anchor λ estimates all within 1.4994–1.5001 of the configured λ=1.5. Full audit in Qwen3.8-Flash-Next-Abliterated.halogen.overlay.hgn.manifest.json.
  2. HGN-HQ conversion. Current upstream flashnext2hgn pipeline: routed experts → weighted Q4CP with the universal expert codebook (x² weighting, no imatrix), dense tensors → q8g32 overlay, MTP sidecar → q8g64, PLE table → FP8 E4M3 + global scale, vision tower → BF16. Few-row dense exceptions stay BF16 per upstream policy.
  3. Direction-preserving expert repair. All 49 fused *.mlp.experts.down_proj.weight tensors (48 trunk + 1 MTP) are patched nibble-by-nibble so that rᵀQ ≈ rᵀW_ablit — the repair only flips adjacent Q4 levels when that reduces directional error. It does not push the projection toward zero, so it does not silently turn λ=1.5 into λ≈1.

Verification (shipped artifacts, measured)

From Qwen3.8-Flash-Next-Abliterated.experts.json:

metric value gate
repaired expert tensors 49 48 trunk + 1 MTP
max direction residual ratio 0.0201 ≤ 0.05 (hard fail)
mean direction residual ratio 0.0200 —
weighted error regressions vs absmax baseline 0 0 (fail-closed restore)
max nibble flip fraction 0.154 % —

Any tensor the repair would have made worse than the same-codebook absmax baseline was restored to the upstream quantized version and the run failed closed; zero tensors hit that path.

No downstream benchmark re-run was performed on the quantized weights — the capability and refusal numbers on the source model card are properties of the bf16 build. What is verified here is that the λ=1.5 direction edit survives quantization to within a 2 % residual on every fused expert tensor.

Files

file size role
Qwen3.8-Flash-Next-Abliterated.hgn 115.5 GiB base weights (routed experts Q4CP, dense base)
Qwen3.8-Flash-Next-Abliterated.overlay.hgn 4.83 GiB production main overlay — q8g32 HGN-HQ dense
Qwen3.8-Flash-Next-Abliterated-mtp.hgn 90 MiB MTP draft head sidecar (q8g64)
Qwen3.8-Flash-Next-Abliterated-vision.hgn 858 MiB vision tower (BF16)
Qwen3.8-Flash-Next-Abliterated.halogen.overlay.hgn 2.31 GiB phase-1 audit artifact — do not load in production
…halogen.overlay.hgn.manifest.json / …report.tsv — direction-recovery audit data
Qwen3.8-Flash-Next-Abliterated.direction.npy 10 KiB recovered refusal direction r, (2560,) float32, ‖r‖≈1
Qwen3.8-Flash-Next-Abliterated.experts.json 730 KiB per-tensor expert repair report
tokenizer/ — tokenizer, chat template, processor configs

Load order matters:

base .hgn  →  .overlay.hgn (q8g32 main)  →  -mtp.hgn  →  -vision.hgn

Do not additionally load .halogen.overlay.hgn after the q8g32 main overlay — it is a legacy calibration-transfer artifact and would overwrite the current 8-bit dense tensors with 4-bit Halogen-derived ones, defeating the HGN-HQ route.

Runtime

  • halogen-flash-server or gfx1151-engine, with HGN-HQ (q8g32 overlay) support.
  • Converted and validated on a Ryzen AI Max+ 395 (gfx1151) with the TheRock / ROCm 10.0.0 stable stack; on that machine the 10.1/10.2 stacks fail a rocBLAS SGEMM M=640,N=1,K=2560 probe, so 10.0.0 is the known-good runtime.
  • On startup the engine should report dense: 8-bit (q8g32) weights present, few-row bf16 gemv on.

License

Qwen Community License 1.0 — inherited from the base model through the abliterated checkpoint; see the source repo's LICENSE. Redistribution and derivative weights are permitted. Two conditions to read before deploying:

  • Products above 100 M MAU or US$20 M monthly revenue must display the model name prominently.
  • Running a Model-as-a-Service or AI Work Assistant business requires a separate license from Qwen before commercial use. Internal use is exempt provided the model, its outputs and its capabilities are not made available to third parties.

Support / 打賞

If these models are useful to you, tips are appreciated — they pay for the GPU time. 如果這些模型對你有幫助,歡迎打賞,用於支應算力成本。

USDT (TRC20) · TPTo32r7vKazpTNaFqfFZ2ztoK1DG88888

Disclaimer

This model will not refuse. It is published for alignment and safety research. You are responsible for your use of it and for complying with applicable law.

Downloads last month
41
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DogEasy/Qwen3.8-Flash-Next-Abliterated-HGNv1

Quantized
(3)
this model