Instructions to use WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF # Run inference directly in the terminal: llama cli -hf WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF # Run inference directly in the terminal: llama cli -hf WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF # Run inference directly in the terminal: ./llama-cli -hf WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF
Use Docker
docker model run hf.co/WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF
- LM Studio
- Jan
- vLLM
How to use WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF
- Ollama
How to use WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF with Ollama:
ollama run hf.co/WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF
- Unsloth Desktop
- Pi
How to use WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF with Docker Model Runner:
docker model run hf.co/WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF
- Lemonade
How to use WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF
Run and chat with the model
lemonade run user.Winery-Qwen3.8-Flash-Next-Cuvee-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "WineryLabs/Winery-Qwen3.8-Flash-Next-Cuvee-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
🍷 Winery Qwen3.8-Flash-Next · Cuvée
One blend of Qwen3.8-Flash-Next (125B MoE, 512 experts, 10 active) in three bottles, sized so it runs on a gaming PC with Winery Strata. That's a fork of Strata with a wine-red glass UI, where Cuvée is the first model in the installer.
- Short thinking from Swift 1.5 (UkisAI)
- Refusals poured out by Huihui's abliteration
- Base everywhere else: routed experts' gate/up, routers, hyper-connections, PLE n-gram table
🍷 Visit the Cuvée Space for the expert-cellar map and a bottle picker for your RAM.
The bottles
| Folder | For | Experts | Expert types | Size |
|---|---|---|---|---|
POCKET/ |
8 GB GPU · 24-32 GB RAM | 256 per layer | IQ2_XS gate/up · Q2_0 down, lean dense (Q4_K attn) | ~48 GB |
SLIM/ |
32 GB RAM | 256 per layer | IQ3_XXS gate/up · IQ4_NL down | ~60 GB |
RESERVE/ |
64 GB RAM | all 512 | IQ3_XXS gate/up · IQ4_NL down | ~86 GB |
GRAND/ |
96-128 GB+ RAM | all 512 | Q4_K gate/up · Q5_1 down | ~113 GB |
All three keep the dense weights in Q6_K/Q8_0 and the 28.8 GB PLE table in IQ4_NL (Strata reads it from the SSD). They were quantised with Unsloth's imatrix for Flash-Next.
Slim drops half the experts, the same move as ISTA-DASLab's Coder. The difference is that it keeps the 256 experts routed most often on general text (Unsloth's calibration routing counts, 72-91% of routed tokens per layer, 77% on average), not the ones code uses. Expect it to be a little weaker than Reserve, mostly on rare topics.
How it was made
Comparing the fine-tunes with the base tensor by tensor showed they touch different parts of the model:
| Tensors | Swift 1.5 | Huihui |
|---|---|---|
| attention q/k/v · DeltaNet qkv/z · shared-expert gate/up | ✓ | |
| attention/DeltaNet output · shared-expert down | ✓ | ✓ |
| routed experts' down | ✓ | |
| everything else |
So Cuvée = base + (Swift − base) + (Huihui − base) (task arithmetic, FP32 per tensor). Every tensor was streamed straight from the three BF16 checkpoints with HTTP range requests, merged and written to a Q8_0 GGUF. The Q8_0 was then quantised once to each bottle. The vision tower and MTP head are the original model's (Strata fetches them itself).
Run it
With Winery Strata (NVIDIA RTX 20-50 or a recent AMD card; 8 GB cards such as the RTX 4060 laptop GPU use POCKET, which setup picks for them by itself):
git clone https://huggingface.co/WineryLabs/Winery-Strata && cd Winery-Strata
START-HERE.bat --setup --family winery --model POCKET # Windows, 8 GB GPU; SLIM / RESERVE / GRAND for more
START-HERE.bat --setup --family winery --model SLIM # Windows; RESERVE / GRAND for more RAM
./setup.sh --setup --family winery --model RESERVE # Linux
With llama.cpp (a build with qwen4exp support), e.g. llama-server -m RESERVE/Winery-Cuvee-RESERVE-00001-of-00002.gguf -c 32768 -ngl 99 --n-cpu-moe 48.
Not measured yet: speed on real hardware, and benchmark scores against the parent models. Sanity chats on CPU are in
TASTING.md.
Licence
A derivative of Qwen3.8-Flash-Next (Qwen Community License 1.0, LICENSE-QWEN) that contains UkisAI's
Swift Contribution (Swift Open License v1.0, LICENSE: free, including commercial use below US$1M yearly
revenue). Huihui's weights are under the Qwen licence. Attribution and our changes are in NOTICE.
- Downloads last month
- 2,005
We're not able to determine the quantization variants.