Qwen3.8-Flash-Next — Abliterated · HGN-HQ (λ=1.5)
HGN-HQ quantized build of
windowsxp811203/Qwen3.8-Flash-Next-Abliterated
for halogen-flash-server and the
gfx1151-engine runtime.
This is not a transformers / vLLM / llama.cpp checkpoint. There are no safetensors or GGUF files
here; the .hgn set only loads in an HGN-compatible engine. For the bf16 weights, use the source
repo above.
What this build is
The source checkpoint removes refusal by orthogonalizing 151 residual-stream-writing tensors against
a captured direction at λ=1.5 (W_ablit = (I − λ r rᵀ) W_stock). Naively quantizing those edited
tensors would blur exactly the edit the model exists for — so this build quantizes around the
direction and then repairs the experts against it:
- Direction recovery. The shared refusal direction
ris recovered from the official Halogen quality overlay (qwen38-flash-next-w4b.overlay.hgn) and the abliterated BF16 checkpoint, using 12 anchor tensors (allself_attn.o_proj/linear_attn.out_proj/shared_expert.down_projspread across the 48 layers). Recovery was effectively exact: min anchor cosine 0.99995, median s1/s2 ≈ 57, per-anchor λ estimates all within 1.4994–1.5001 of the configured λ=1.5. Full audit inQwen3.8-Flash-Next-Abliterated.halogen.overlay.hgn.manifest.json. - HGN-HQ conversion. Current upstream
flashnext2hgnpipeline: routed experts → weighted Q4CP with theuniversalexpert codebook (x² weighting, no imatrix), dense tensors → q8g32 overlay, MTP sidecar → q8g64, PLE table → FP8 E4M3 + global scale, vision tower → BF16. Few-row dense exceptions stay BF16 per upstream policy. - Direction-preserving expert repair. All 49 fused
*.mlp.experts.down_proj.weighttensors (48 trunk + 1 MTP) are patched nibble-by-nibble so thatrᵀQ ≈ rᵀW_ablit— the repair only flips adjacent Q4 levels when that reduces directional error. It does not push the projection toward zero, so it does not silently turn λ=1.5 into λ≈1.
Verification (shipped artifacts, measured)
From Qwen3.8-Flash-Next-Abliterated.experts.json:
| metric | value | gate |
|---|---|---|
| repaired expert tensors | 49 | 48 trunk + 1 MTP |
| max direction residual ratio | 0.0201 | ≤ 0.05 (hard fail) |
| mean direction residual ratio | 0.0200 | — |
| weighted error regressions vs absmax baseline | 0 | 0 (fail-closed restore) |
| max nibble flip fraction | 0.154 % | — |
Any tensor the repair would have made worse than the same-codebook absmax baseline was restored to the upstream quantized version and the run failed closed; zero tensors hit that path.
No downstream benchmark re-run was performed on the quantized weights — the capability and refusal numbers on the source model card are properties of the bf16 build. What is verified here is that the λ=1.5 direction edit survives quantization to within a 2 % residual on every fused expert tensor.
Files
| file | size | role |
|---|---|---|
Qwen3.8-Flash-Next-Abliterated.hgn |
115.5 GiB | base weights (routed experts Q4CP, dense base) |
Qwen3.8-Flash-Next-Abliterated.overlay.hgn |
4.83 GiB | production main overlay — q8g32 HGN-HQ dense |
Qwen3.8-Flash-Next-Abliterated-mtp.hgn |
90 MiB | MTP draft head sidecar (q8g64) |
Qwen3.8-Flash-Next-Abliterated-vision.hgn |
858 MiB | vision tower (BF16) |
Qwen3.8-Flash-Next-Abliterated.halogen.overlay.hgn |
2.31 GiB | phase-1 audit artifact — do not load in production |
…halogen.overlay.hgn.manifest.json / …report.tsv |
— | direction-recovery audit data |
Qwen3.8-Flash-Next-Abliterated.direction.npy |
10 KiB | recovered refusal direction r, (2560,) float32, ‖r‖≈1 |
Qwen3.8-Flash-Next-Abliterated.experts.json |
730 KiB | per-tensor expert repair report |
tokenizer/ |
— | tokenizer, chat template, processor configs |
Load order matters:
base .hgn → .overlay.hgn (q8g32 main) → -mtp.hgn → -vision.hgn
Do not additionally load .halogen.overlay.hgn after the q8g32 main overlay — it is a legacy
calibration-transfer artifact and would overwrite the current 8-bit dense tensors with 4-bit
Halogen-derived ones, defeating the HGN-HQ route.
Runtime
halogen-flash-serverorgfx1151-engine, with HGN-HQ (q8g32 overlay) support.- Converted and validated on a Ryzen AI Max+ 395 (gfx1151) with the TheRock /
ROCm 10.0.0 stable stack; on that machine the 10.1/10.2 stacks fail a rocBLAS SGEMM
M=640,N=1,K=2560probe, so 10.0.0 is the known-good runtime. - On startup the engine should report
dense: 8-bit (q8g32) weights present, few-row bf16 gemv on.
License
Qwen Community License 1.0 — inherited from the base model through the abliterated checkpoint;
see the source repo's LICENSE. Redistribution and derivative weights are permitted. Two conditions
to read before deploying:
- Products above 100 M MAU or US$20 M monthly revenue must display the model name prominently.
- Running a Model-as-a-Service or AI Work Assistant business requires a separate license from Qwen before commercial use. Internal use is exempt provided the model, its outputs and its capabilities are not made available to third parties.
Support / 打賞
If these models are useful to you, tips are appreciated — they pay for the GPU time. 如果這些模型對你有幫助,歡迎打賞,用於支應算力成本。
USDT (TRC20) · TPTo32r7vKazpTNaFqfFZ2ztoK1DG88888
Disclaimer
This model will not refuse. It is published for alignment and safety research. You are responsible for your use of it and for complying with applicable law.
- Downloads last month
- 41
Model tree for DogEasy/Qwen3.8-Flash-Next-Abliterated-HGNv1
Base model
Qwen/Qwen3.8-Flash-Next