How to use from
OpenClaw
Start the llama.cpp server
# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama serve -hf WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF
Configure OpenClaw
# Install OpenClaw:
npm install -g openclaw@latest
# Register the local server and set it as the default model:
openclaw onboard --non-interactive --mode local \
  --auth-choice custom-api-key \
  --custom-base-url http://127.0.0.1:8080/v1 \
  --custom-model-id "WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF" \
  --custom-provider-id llama-cpp \
  --custom-compatibility openai \
  --custom-text-input \
  --accept-risk \
  --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quick Links

Winery Qwen3.8-Flash-Next Cuvée

🍷 Winery Qwen3.8-Flash-Next · Cuvée

One blend of Qwen3.8-Flash-Next (125B MoE, 512 experts, 10 active) in three bottles, sized so it runs on a gaming PC with Winery Strata. That's a fork of Strata with a wine-red glass UI, where Cuvée is the first model in the installer.

  • Short thinking from Swift 1.5 (UkisAI)
  • Refusals poured out by Huihui's abliteration
  • Base everywhere else: routed experts' gate/up, routers, hyper-connections, PLE n-gram table

🍷 Visit the Cuvée Space for the expert-cellar map and a bottle picker for your RAM.

The bottles

Folder For Experts Expert types Size
POCKET/ 8 GB GPU · 24-32 GB RAM 256 per layer IQ2_XS gate/up · Q2_0 down, lean dense (Q4_K attn) ~48 GB
SLIM/ 32 GB RAM 256 per layer IQ3_XXS gate/up · IQ4_NL down ~60 GB
RESERVE/ 64 GB RAM all 512 IQ3_XXS gate/up · IQ4_NL down ~86 GB
GRAND/ 96-128 GB+ RAM all 512 Q4_K gate/up · Q5_1 down ~113 GB

All three keep the dense weights in Q6_K/Q8_0 and the 28.8 GB PLE table in IQ4_NL (Strata reads it from the SSD). They were quantised with Unsloth's imatrix for Flash-Next.

Slim drops half the experts, the same move as ISTA-DASLab's Coder. The difference is that it keeps the 256 experts routed most often on general text (Unsloth's calibration routing counts, 72-91% of routed tokens per layer, 77% on average), not the ones code uses. Expect it to be a little weaker than Reserve, mostly on rare topics.

How it was made

Comparing the fine-tunes with the base tensor by tensor showed they touch different parts of the model:

Tensors Swift 1.5 Huihui
attention q/k/v · DeltaNet qkv/z · shared-expert gate/up ✓
attention/DeltaNet output · shared-expert down ✓ ✓
routed experts' down ✓
everything else

So Cuvée = base + (Swift − base) + (Huihui − base) (task arithmetic, FP32 per tensor). Every tensor was streamed straight from the three BF16 checkpoints with HTTP range requests, merged and written to a Q8_0 GGUF. The Q8_0 was then quantised once to each bottle. The vision tower and MTP head are the original model's (Strata fetches them itself).

Run it

With Winery Strata (NVIDIA RTX 20-50 or a recent AMD card; 8 GB cards such as the RTX 4060 laptop GPU use POCKET, which setup picks for them by itself):

git clone https://huggingface.co/WineryLabs/Winery-Strata && cd Winery-Strata
START-HERE.bat --setup --family winery --model POCKET      # Windows, 8 GB GPU; SLIM / RESERVE / GRAND for more
START-HERE.bat --setup --family winery --model SLIM        # Windows; RESERVE / GRAND for more RAM
./setup.sh --setup --family winery --model RESERVE          # Linux

With llama.cpp (a build with qwen4exp support), e.g. llama-server -m RESERVE/Winery-Cuvee-RESERVE-00001-of-00002.gguf -c 32768 -ngl 99 --n-cpu-moe 48.

Not measured yet: speed on real hardware, and benchmark scores against the parent models. Sanity chats on CPU are in TASTING.md.

Licence

A derivative of Qwen3.8-Flash-Next (Qwen Community License 1.0, LICENSE-QWEN) that contains UkisAI's Swift Contribution (Swift Open License v1.0, LICENSE: free, including commercial use below US$1M yearly revenue). Huihui's weights are under the Qwen licence. Attribution and our changes are in NOTICE.

Downloads last month
2,005
GGUF
Model size
177B params
Architecture
qwen4exp
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF

Spaces using WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF 2