Instructions to use nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF:Q4_K_M
Use Docker
docker model run hf.co/nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF with Ollama:
ollama run hf.co/nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF with Docker Model Runner:
docker model run hf.co/nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF:Q4_K_M
- Lemonade
How to use nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Qwen3.8-27B EfficientThink GGUF
- Measured quantization-tier overview
- Lynn Agent v0.87.0
- llama.cpp prerequisites and launch settings
- Why DFlash2, and bundled MTP alternatives
- Per-tier DFlash2 concurrency performance
- Training method
- Published GGUF files
- LynnStyle mixed-precision quantization
- Q6_K
- ไธญๆ่ฏดๆ
- ้ๅๆกฃไฝๅฎๆตๆป่ง
- Lynn Agent v0.87.0
- llama.cpp ็ๆฌไธๅฏๅจ่ฎพ็ฝฎ
- ไธบไปไน้ๆฉ DFlash2๏ผไปฅๅ้ๅ ๆไพ็ MTP ๆนๆก
- ๅๆกฃ DFlash2 ๅนถๅๆง่ฝ
- ่ฎญ็ปๆนๆณ
- ๅทฒๅๅธ GGUF ๆไปถ
- LynnStyle ๆททๅ็ฒพๅบฆ้ๅ
- Q6_K
- Measured quantization-tier overview
Qwen3.8-27B EfficientThink GGUF
No strict loops were observed in the reviewed Q2โQ8 evaluations.
BF16 / FP8 main repository: nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2
Measured quantization-tier overview
Capability scores are shown only after a formal full-suite freeze, with every non-passing sample retained in the denominator. Q8_0, Q6_K, Q5-LynnStyle, Q4-LynnStyle, Q3-LynnStyle, and Q2-LynnStyle are published and have passed four-repository remote path, byte-size, and LFS-SHA256 readback.
| Tier | Repository state | GPQA 198 | MMLU 500 | LCB 100 |
|---|---|---|---|---|
| Q8_0 | Published and remotely verified | 164/198 (82.83%) | 447/500 (89.40%) | 74/100 (74.00%) |
| Q6_K | Published and remotely verified | 171/198 (86.36%) | 440/500 (88.00%) | 78/100 (78.00%) |
| Q5-LynnStyle | Published and remotely verified | 164/198 (82.83%) | 438/500 (87.60%) | 75/100 (75.00%) |
| Q4-LynnStyle | Published and remotely verified | 166/198 (83.84%) | 443/500 (88.60%) | 74/100 (74.00%) |
| Q3-LynnStyle | Published and remotely verified | 172/198 (86.87%) | 435/500 (87.00%) | 78/100 (78.00%) |
| Q2-LynnStyle | Published and remotely verified | 167/198 (84.34%) | 416/500 (83.20%) | 75/100 (75.00%) |
Lynn Agent v0.87.0
Lynn Agent v0.87.0 uses this release's Q2-LynnStyle / Q3-LynnStyle + DFlash2 packages. The pairing passed runtime validation on DGX Spark; notarized Mac Apple Silicon and Intel builds, the Windows installer runtime check, CI in both repositories, matching main heads and tags across all three release repositories, complete SHA256 verification for 23 public files, and a remote CLI installation all passed.
Natural moving tree shadows and soft window light, enabled by default. Hover over Shadows for the off switch location, or click to open Settings. Playback pauses in the background and stays still with reduced motion. Images now participate in file filtering; slash templates replace persistent task mode; translation moved into the message menu; Expert Roundtable is now an optional plugin; and session edit targeting and stop preprocessing were fixed. Kimi Datasource remains available under MCP and requires users to scan the QR code and sign in with their own account.
This client update did not change any model weight, quantized artifact, benchmark score, or performance metric in this repository.
| Installer | China mirror | GitHub fallback |
|---|---|---|
| Mac Apple Silicon | Download | Download |
| Mac Intel | Download | Download |
| Windows | Download | Download |
Release records: primary GitHub repository ยท legacy GitHub repository ยท Gitee ยท CLI package
llama.cpp prerequisites and launch settings
Use an official llama.cpp build containing DFlash2 support, merged 2026-08-27, commit b10f9ca58c89 or a compatible newer build. A reproducible minimum build from the official repository is:
git clone https://github.com/ggml-org/llama.cpp.git llama.cpp-efficientthink
git -C llama.cpp-efficientthink checkout b10f9ca58c89ccfc3653ac01e979dd085d582b76
cmake -S llama.cpp-efficientthink -B llama.cpp-efficientthink/build -DGGML_CUDA=ON
cmake --build llama.cpp-efficientthink/build --config Release -j 4 --target llama-server
export PATH="$PWD/llama.cpp-efficientthink/build/bin:$PATH"
llama-server --version
llama-server --help | grep -- draft-dflash
Run the tier commands below from this repository's download root. If using the main repository's mirror, first cd GGUF. Choose one tier/server at a time. Each command explicitly selects GPU 0; remove or adjust CUDA_VISIBLE_DEVICES=0 only for your own multi-GPU configuration. CUDA Toolkit and a compiler supported by the chosen GPU are prerequisites; the Spark FP8 Docker image is not required for GGUF.
The measured layout is 8 slots ร 32,768 tokens = --ctx-size 262144 total, not 32K total shared by eight slots. C4/C8 means four/eight active client requests against the eight-slot server. On smaller-memory hardware reduce both slots and total context (for example one slot with --parallel 1 --ctx-size 32768); those settings do not inherit the published C4/C8 throughput. Weight size alone is not the total runtime memory requirement.
DFlash2 can use either the separate Q8_0 or Q4_K_M draft, --spec-type draft-dflash, --spec-draft-n-max 7, and --spec-draft-n-min 0. It is not the in-model MTP path. --jinja and explicit XH template arguments retain the model's thinking mode. The examples bind to 127.0.0.1:8080; expose a different interface only with appropriate access controls. The --mmproj sidecar is for images; omit that line for text-only serving.
Official references: server options, speculative decoding, CUDA build. These corrected commands match the documented flags and measured slot layout; the build instructions are not a new six-tier benchmark run.
Why DFlash2, and bundled MTP alternatives
DFlash2 remains our preferred speculative decoder for its measured draft acceptance and output TPS. Parallel block drafting and acceptance of multiple tokens per verification help reduce sequential decoding overhead. Method reference: DFlash authors.
For example, the measured Q8_0 C4 configuration achieved 60.98% acceptance and 246 aggregate tok/s; Q3-LynnStyle C4 achieved 58.00% and 269 aggregate tok/s. The tables below report each tier independently: C4 is the practical balance and C8 the maximum measured throughput. Acceptance and TPS depend on the model, hardware, workload, and concurrency; these results are not a matched MTP comparison.
Q4_0 and Q8_0 MTP GGUF files are bundled
Every Q2-LynnStyle through Q8_0 directory now contains the same two independently verified MTP sidecars:
| File | Role | Size | SHA256 |
|---|---|---|---|
mtp-Qwen3.8-27B-Q4_0.gguf |
Smaller Q4_0 MTP draft | 1,680,271,648 bytes | 051a1764cff8c4f3ee6ae8b00593a0364c7539c67fa50ffc58f3f96509fca38e |
mtp-Qwen3.8-27B-Q8_0.gguf |
Higher-precision Q8_0 MTP draft | 3,164,006,688 bytes | cbf60a0c48b431bb61f1d49b8948dc88ac29c398d6dbdbbb2e6e89ef77eacc9a |
Both are derived from ggml-org/Qwen3.8-27B-GGUF@0669b98607d47046c7c2b3f801011d54a08cfccf and parse as Qwen MTP GGUFs with 65 blocks and 18 tensors. A real DGX Spark smoke paired each file with the published Q3-LynnStyle main model: both loaded, returned HTTP 200, generated non-empty final content, and stopped normally. This is a compatibility smoke, not a frozen MTP TPS or capability benchmark.
Choose one speculative path per launch. Do not combine DFlash2 and MTP. Example with the smaller MTP draft:
CUDA_VISIBLE_DEVICES=0 llama-server \
--model Q3-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q3-LynnStyle.gguf \
--model-draft Q3-LynnStyle/mtp-Qwen3.8-27B-Q4_0.gguf \
--spec-type draft-mtp \
--gpu-layers all --gpu-layers-draft all \
--parallel 1 --ctx-size 32768 --flash-attn on \
--jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
--host 127.0.0.1 --port 8080
Substitute mtp-Qwen3.8-27B-Q8_0.gguf for the Q8_0 MTP option. Add the tier's --mmproj only when serving images.
Per-tier DFlash2 concurrency performance
Common measured protocol: one GPU for the main model, DFlash2, --parallel 8 --ctx-size 262144 (8 slots ร 32K), C1/C2/C4/C8, fixed 256-token pressure outputs, and a 60-second request ceiling. DFlash2 used n_max=7 and n_min=0. Throughput is aggregate completion tok/s.
Q8_0
| Concurrency | Aggregate throughput | DFlash2 acceptance | Mean accepted length / 8 |
|---|---|---|---|
| C1 | 101.46 tok/s | 51.30% | 4.54 |
| C2 | 165.63 tok/s | 54.47% | 4.76 |
| C4 | 245.94 tok/s | 60.98% | 5.25 |
| C8 | 255.71 tok/s | 52.01% | 4.60 |
- Concurrency guidance: C4 practical balance; C8 maximum throughput.
- 15 fixed-length requests per tier; zero timeout, request error, empty output, or punctuation collapse.
Q6_K
| Concurrency | Aggregate throughput | DFlash2 acceptance | Mean accepted length / 8 |
|---|---|---|---|
| C1 | 91.43 tok/s | 40.83% | 3.79 |
| C2 | 165.08 tok/s | 48.28% | 4.36 |
| C4 | 242.11 tok/s | 49.00% | 4.37 |
| C8 | 279.83 tok/s | 46.80% | 4.24 |
- Concurrency guidance: C4 lower-latency balance; C8 maximum throughput.
- Zero timeout, request error, empty output, or punctuation collapse across the sweep.
Q5-LynnStyle
| Concurrency | Aggregate throughput | DFlash2 acceptance | Mean accepted length / 8 |
|---|---|---|---|
| C1 | 94.64 tok/s | 40.87% | 3.81 |
| C2 | 158.71 tok/s | 39.98% | 3.78 |
| C4 | 238.06 tok/s | 48.22% | 4.35 |
| C8 | 275.73 tok/s | 48.38% | 4.34 |
- Concurrency guidance: C4 practical balance; C8 maximum throughput.
- Zero timeout, request error, empty output, or punctuation collapse across the sweep.
Q4-LynnStyle
| Concurrency | Aggregate throughput | DFlash2 acceptance | Mean accepted length / 8 |
|---|---|---|---|
| C1 | 97.96 tok/s | 38.83% | 3.70 |
| C2 | 149.27 tok/s | 36.15% | 3.51 |
| C4 | 256.52 tok/s | 47.69% | 4.30 |
| C8 | 263.50 tok/s | 42.88% | 3.95 |
- Concurrency guidance: C4 practical balance; C8 maximum throughput.
- Zero timeout, request error, empty output, or punctuation collapse across the sweep.
Q3-LynnStyle
| Concurrency | Aggregate throughput | DFlash2 acceptance | Mean accepted length / 8 |
|---|---|---|---|
| C1 | 112.97 tok/s | 47.22% | 4.25 |
| C2 | 210.78 tok/s | 57.22% | 5.00 |
| C4 | 269.32 tok/s | 58.00% | 5.02 |
| C8 | 286.16 tok/s | 52.41% | 4.63 |
- Concurrency guidance: C4 is the practical balance; C8 is the maximum measured throughput.
- The full sweep had zero timeout, request error, empty output, or punctuation collapse.
Training method
Qwen/Qwen3.8-27B โ capability-preserving SFT โ terminal-behavior SimPO โ per-tensor FP32 delta merge โ BF16 โ GGUF quantization.
- SFT: 1,905 samples, 1 epoch, 239 optimizer steps, effective batch 8; LoRA r=16, alpha=32, dropout=0.05; LR 5e-6, 12 warmup steps, seed 20260901.
- SimPO: 110 preference pairs over 73 unique prompts, 5 optimizer steps; beta=1.0, gamma=0.2, peak LR 5e-7; LoRA r=16, alpha=32, dropout=0, seed 20260903, world size 2, FSDP full sharding.
- Training hardware: 2x NVIDIA RTX PRO 6000 Blackwell Server Edition.
This repository publishes text-main GGUF files with independently verified vision and DFlash2 sidecars. Q8_0 and Q6_K use the Q8 vision projector; Q5-LynnStyle, Q4-LynnStyle, and Q3-LynnStyle use the Q6 vision projector. All six tiers provide Q8_0 and smaller Q4_K_M DFlash2 drafts plus Q4_0 and Q8_0 MTP sidecars. Choose exactly one speculative path per launch; none of these sidecars is a main model.
Published GGUF files
Q8_0
Formal full-suite evaluation: GPQA 164/198 (82.83%) ยท MMLU 447/500 (89.40%) ยท LCB 74/100 (74.00%).
| Path | Role | Size |
|---|---|---|
Q8_0/Qwen3.8-27B-EfficientThink-SimPO-Q8_0.gguf |
Q8_0 text-main model | 28,595,763,552 bytes |
Q8_0/mmproj-Qwen3.8-27B-Q8_0.gguf |
Q8 vision projector | 629,247,008 bytes |
Q8_0/dflash2-qwen38-27b-Q8_0.gguf |
Q8_0 DFlash2 draft | 2,056,414,720 bytes |
Q8_0/dflash2-qwen38-27b-Q4_K_M.gguf |
Optional Q4_K_M DFlash2 draft; not MTP | 1,143,006,720 bytes |
Q8_0/mtp-Qwen3.8-27B-Q4_0.gguf |
Optional Q4_0 MTP draft; not a main model or DFlash2 | 1,680,271,648 bytes |
Q8_0/mtp-Qwen3.8-27B-Q8_0.gguf |
Optional Q8_0 MTP draft; not a main model or DFlash2 | 3,164,006,688 bytes |
Q8_0/SHA256SUMS |
Checksums | 396 bytes |
Q8_0/manifest.json |
Structured manifest | 4,621 bytes |
Concurrency guidance: C4 is the practical balance (246 tok/s, 60.98% acceptance); C8 is the maximum measured throughput (256 tok/s, 52.01%).
CUDA_VISIBLE_DEVICES=0 llama-server \
--model Q8_0/Qwen3.8-27B-EfficientThink-SimPO-Q8_0.gguf \
--mmproj Q8_0/mmproj-Qwen3.8-27B-Q8_0.gguf \
--model-draft Q8_0/dflash2-qwen38-27b-Q8_0.gguf \
--spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
--gpu-layers all --gpu-layers-draft all \
--parallel 8 --ctx-size 262144 --flash-attn on \
--jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
--host 127.0.0.1 --port 8080
Q5-LynnStyle
Formal full-suite evaluation: GPQA 164/198 (82.83%) ยท MMLU 438/500 (87.60%) ยท LCB 75/100 (75.00%). LCB: Easy 23/23, Medium 29/31, Hard 23/46.
| Path | Role | Size |
|---|---|---|
Q5-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q5-LynnStyle.gguf |
Q5-LynnStyle text-main model | 22,412,452,608 bytes |
Q5-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf |
Q6 vision projector | 618,390,560 bytes |
Q5-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf |
Q8_0 DFlash2 draft | 2,056,414,720 bytes |
Q5-LynnStyle/dflash2-qwen38-27b-Q4_K_M.gguf |
Optional Q4_K_M DFlash2 draft; not MTP | 1,143,006,720 bytes |
Q5-LynnStyle/mtp-Qwen3.8-27B-Q4_0.gguf |
Optional Q4_0 MTP draft; not a main model or DFlash2 | 1,680,271,648 bytes |
Q5-LynnStyle/mtp-Qwen3.8-27B-Q8_0.gguf |
Optional Q8_0 MTP draft; not a main model or DFlash2 | 3,164,006,688 bytes |
Q5-LynnStyle/SHA256SUMS |
Checksums | 501 bytes |
Q5-LynnStyle/manifest.json |
Structured manifest | 5,287 bytes |
Concurrency guidance: C4 is the practical balance (238 tok/s, 48.22% acceptance); C8 is the maximum measured throughput (276 tok/s, 48.38%).
CUDA_VISIBLE_DEVICES=0 llama-server \
--model Q5-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q5-LynnStyle.gguf \
--mmproj Q5-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf \
--model-draft Q5-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf \
--spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
--gpu-layers all --gpu-layers-draft all \
--parallel 8 --ctx-size 262144 --flash-attn on \
--jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
--host 127.0.0.1 --port 8080
Q4-LynnStyle
Formal full-suite evaluation: GPQA 166/198 (83.84%) ยท MMLU 443/500 (88.60%) ยท LCB 74/100 (74.00%). LCB: Easy 23/23, Medium 30/31, Hard 21/46; 100 unique tasks, errors=0, clean gate PASS.
| Path | Role | Size |
|---|---|---|
Q4-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q4-LynnStyle.gguf |
Q4-LynnStyle text-main model | 19,351,954,688 bytes |
Q4-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf |
Q6 vision projector | 618,390,560 bytes |
Q4-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf |
Shared Q8_0 DFlash2 draft | 2,056,414,720 bytes |
Q4-LynnStyle/dflash2-qwen38-27b-Q4_K_M.gguf |
Optional Q4_K_M DFlash2 draft; not MTP | 1,143,006,720 bytes |
Q4-LynnStyle/mtp-Qwen3.8-27B-Q4_0.gguf |
Optional Q4_0 MTP draft; not a main model or DFlash2 | 1,680,271,648 bytes |
Q4-LynnStyle/mtp-Qwen3.8-27B-Q8_0.gguf |
Optional Q8_0 MTP draft; not a main model or DFlash2 | 3,164,006,688 bytes |
Q4-LynnStyle/SHA256SUMS |
Checksums | 501 bytes |
Q4-LynnStyle/manifest.json |
Structured manifest | 5,439 bytes |
| Concurrency | Aggregate throughput | DFlash2 acceptance | Mean accepted length / 8 |
|---|---|---|---|
| C1 | 97.96 tok/s | 38.83% | 3.70 |
| C2 | 149.27 tok/s | 36.15% | 3.51 |
| C4 | 256.52 tok/s | 47.69% | 4.30 |
| C8 | 263.50 tok/s | 42.88% | 3.95 |
Concurrency guidance: C4 is the practical balance (257 tok/s, 47.69% acceptance); C8 is the maximum measured throughput (264 tok/s, 42.88%). The full sweep had zero timeout, request error, empty output, or punctuation collapse.
CUDA_VISIBLE_DEVICES=0 llama-server \
--model Q4-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q4-LynnStyle.gguf \
--mmproj Q4-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf \
--model-draft Q4-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf \
--spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
--gpu-layers all --gpu-layers-draft all \
--parallel 8 --ctx-size 262144 --flash-attn on \
--jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
--host 127.0.0.1 --port 8080
Q3-LynnStyle
Formal full-suite evaluation: GPQA 172/198 (86.87%) ยท MMLU 435/500 (87.00%) ยท LCB 78/100 (78.00%). GPQA used 198 unique items with errors=0 and strict loop=0; MMLU had errors=0 and parse fail=0; LCB: Easy 23/23, Medium 28/31, Hard 27/46, with 100 unique tasks and errors=0. Every non-passing sample remains in the denominator.
| Path | Role | Size |
|---|---|---|
Q3-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q3-LynnStyle.gguf |
Q3-LynnStyle text-main model | 17,032,184,960 bytes |
Q3-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf |
Q6 vision projector | 618,390,560 bytes |
Q3-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf |
Shared Q8_0 DFlash2 draft | 2,056,414,720 bytes |
Q3-LynnStyle/dflash2-qwen38-27b-Q4_K_M.gguf |
Optional Q4_K_M DFlash2 draft; not MTP | 1,143,006,720 bytes |
Q3-LynnStyle/mtp-Qwen3.8-27B-Q4_0.gguf |
Optional Q4_0 MTP draft; not a main model or DFlash2 | 1,680,271,648 bytes |
Q3-LynnStyle/mtp-Qwen3.8-27B-Q8_0.gguf |
Optional Q8_0 MTP draft; not a main model or DFlash2 | 3,164,006,688 bytes |
Q3-LynnStyle/SHA256SUMS |
Checksums | 501 bytes |
Q3-LynnStyle/manifest.json |
Structured manifest | 5,526 bytes |
Concurrency guidance: C4 is the practical balance (269 tok/s, 58.00% acceptance); C8 is the maximum measured throughput (286 tok/s, 52.41%). The full sweep had zero timeout, request error, empty output, or punctuation collapse.
CUDA_VISIBLE_DEVICES=0 llama-server \
--model Q3-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q3-LynnStyle.gguf \
--mmproj Q3-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf \
--model-draft Q3-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf \
--spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
--gpu-layers all --gpu-layers-draft all \
--parallel 8 --ctx-size 262144 --flash-attn on \
--jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
--host 127.0.0.1 --port 8080
Q2-LynnStyle
Formal full-suite scores: GPQA 167/198 (84.34%) ยท MMLU 416/500 (83.20%) ยท LCB 75/100 (75.00%). LCB: Easy 23/23, Medium 29/31, Hard 23/46; all 23 length-limited 32K outputs remain failures in the denominator.
The final LCB view combines the byte-identical 99 unaffected rows from the original run with one Lynn-authorized exact retry of atcoder:arc181_b:81. The original request ended in a streaming JSON error. The retry used the same model, runner, request, and parameters and received a complete DONE, but again reached 32K with no submitted code, so the score remains 75/100. This is not represented as one uninterrupted error-free run.
| Path | Role | Size |
|---|---|---|
Q2-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q2-LynnStyle.gguf |
Q2-LynnStyle text-main model | 12,999,977,600 bytes |
Q2-LynnStyle/mmproj-Qwen3.8-27B-Q4_K_M.gguf |
Q4_K_M vision projector | 522,293,024 bytes |
Q2-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf |
Optional Q8_0 DFlash2 draft; not MTP | 2,056,414,720 bytes |
Q2-LynnStyle/dflash2-qwen38-27b-Q4_K_M.gguf |
Optional Q4_K_M DFlash2 draft; not MTP | 1,143,006,720 bytes |
Q2-LynnStyle/mtp-Qwen3.8-27B-Q4_0.gguf |
Optional Q4_0 MTP draft; not a main model or DFlash2 | 1,680,271,648 bytes |
Q2-LynnStyle/mtp-Qwen3.8-27B-Q8_0.gguf |
Optional Q8_0 MTP draft; not a main model or DFlash2 | 3,164,006,688 bytes |
Q2-LynnStyle/SHA256SUMS |
Checksums | 406 bytes |
Q2-LynnStyle/manifest.json |
Structured manifest | 5,504 bytes |
Q2-LynnStyle uses GSQ-RCO mixed-precision quantization with IQ numerical refinement over 70 selected MLP matrices; attention, head, and SSM tensors were not numerically refined. If memory permits, Q3-LynnStyle or a higher tier remains the safer quality choice. This exact 12,999,977,600-byte build has no frozen TPS result, so throughput from the retired engineering Q2 is not reused.
CUDA_VISIBLE_DEVICES=0 llama-server \
--model Q2-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q2-LynnStyle.gguf \
--mmproj Q2-LynnStyle/mmproj-Qwen3.8-27B-Q4_K_M.gguf \
--model-draft Q2-LynnStyle/dflash2-qwen38-27b-Q4_K_M.gguf \
--spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
--gpu-layers all --gpu-layers-draft all \
--parallel 4 --ctx-size 131072 --flash-attn on \
--jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
--host 127.0.0.1 --port 8080
The Q8_0, Q6_K, Q5-LynnStyle, Q4-LynnStyle, Q3-LynnStyle, and Q2-LynnStyle main GGUF files are GGUF v3, qwen35, 64-layer, 851-tensor text trunks with no embedded vision or MTP tensors; all passed remote byte/SHA256 readback. The published Q8_0 DFlash2 passed static parsing and a representative Q2-LynnStyle runtime smoke. Full per-tier concurrency retests remain pending, and historical DFlash2 grids are not attributed to this Q8_0 file.
LynnStyle mixed-precision quantization
Q5/Q4/Q3 LynnStyle tiers are quantized directly from the final SimPO F16 text trunk with the frozen 512-chunk imatrix (496 entries), without requantization. Protection combines T10 REAP, first/last-layer and SSM guards, plus per-tensor full-attention, SSM, and FFN audits; higher precision wins when rules overlap.
- Q5-LynnStyle (published): Q5 base; Q8 for L0โ3, L60โ63, and critical REAP layers 50/52/54/56/58; Q6 for high-sensitivity layers 21/42/44/46/48/49/53/57/59, full-attention q/k/v/o, and selected middle
ssm_out. Of 498 converted tensors: Q8_0 104, Q6_K 186, Q5_K 208. - Q4-LynnStyle (published): Q4 base; Q8 for L0โ3 and L60โ63; Q6 for critical shoulder layers 50/52/54/56โ59; Q5 for high layers 21/42/44/46/48/49/53; Q6 for full-attention q/k/v/o and selected
ssm_out. Distribution: Q8 62, Q6 156, Q5 172, Q4 108. - Q3-LynnStyle (published): Q3 base; Q6 for L0โ3 and L60โ63; Q5 for L4โ6 and 50/52/54/56โ59; Q4 for high layers 21/42/44/46/48/49/53; Q5 for full-attention and middle
ssm_out. Distribution: Q8 2, Q6 62, Q5 157, Q4 199, Q3 78.
Hub sidebars may expose only nominal GGUF general.file_type, which cannot fully describe custom per-tensor mixed precision. Treat each tier's manifest.json, SHA256SUMS, and actual GGUF parse as authoritative.
Q6_K
Formal full-suite evaluation: GPQA 171/198 (86.36%) ยท MMLU 440/500 (88.00%) ยท LCB 78/100 (78.00%).
LCB difficulty split: Easy 23/23, Medium 30/31, Hard 25/46.
| Concurrency | Aggregate throughput | DFlash2 acceptance | Mean accepted length / 8 | Maximum request time |
|---|---|---|---|---|
| C1 | 91 tok/s | 40.83% | 3.79 | 2.80s |
| C2 | 165 tok/s | 48.28% | 4.36 | 3.10s |
| C4 (lower-latency balance) | 242 tok/s | 49.00% | 4.37 | 4.23s |
| C8 (maximum throughput) | 280 tok/s | 46.80% | 4.24 | 7.31s |
Request errors, timeouts, empty outputs, and punctuation collapse were zero at all four tested concurrency levels. This serving sweep characterizes throughput and speculative acceptance; it does not replace formal quality evaluation.
Files
| Path | Role | Size |
|---|---|---|
Q6_K/Qwen3.8-27B-EfficientThink-SimPO-Q6_K.gguf |
Q6_K text-main model | 22,082,530,048 bytes |
Q6_K/mmproj-Qwen3.8-27B-Q8_0.gguf |
Q8 vision projector | 629,247,008 bytes |
Q6_K/dflash2-qwen38-27b-Q8_0.gguf |
Q8_0 DFlash2 draft | 2,056,414,720 bytes |
Q6_K/dflash2-qwen38-27b-Q4_K_M.gguf |
Optional Q4_K_M DFlash2 draft; not MTP | 1,143,006,720 bytes |
Q6_K/mtp-Qwen3.8-27B-Q4_0.gguf |
Optional Q4_0 MTP draft; not a main model or DFlash2 | 1,680,271,648 bytes |
Q6_K/mtp-Qwen3.8-27B-Q8_0.gguf |
Optional Q8_0 MTP draft; not a main model or DFlash2 | 3,164,006,688 bytes |
Recommended launch
CUDA_VISIBLE_DEVICES=0 llama-server \
--model Q6_K/Qwen3.8-27B-EfficientThink-SimPO-Q6_K.gguf \
--mmproj Q6_K/mmproj-Qwen3.8-27B-Q8_0.gguf \
--model-draft Q6_K/dflash2-qwen38-27b-Q8_0.gguf \
--spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
--gpu-layers all --gpu-layers-draft all \
--parallel 8 --ctx-size 262144 --flash-attn on \
--jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
--host 127.0.0.1 --port 8080
ไธญๆ่ฏดๆ
ๅทฒๅฎกๆฅ็ Q2โQ8 ๆต่ฏไธญๆชๅ็ฐไธฅๆ ผๆญปๅพช็ฏใ
BF16 / FP8 ไธปไป๏ผnerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2
้ๅๆกฃไฝๅฎๆตๆป่ง
่ฝๅๅๆฐๅชๅๆญฃๅผๅ จ้ๅป็ป็ปๆ๏ผๆๆๆช้่ฟๆ ทๆฌๅไฟ็ๅจๅๆฏใQ8_0ใQ6_KใQ5-LynnStyleใQ4-LynnStyleใQ3-LynnStyle ไธ Q2-LynnStyle ๅๅทฒๅๅธ๏ผๅนถๅฎๆๅไป่ฟ็ซฏ่ทฏๅพใๅญ่ๅ LFS SHA256 ๅ่ฏปใ
| ๆกฃไฝ | ไปๅบ็ถๆ | GPQA 198 | MMLU 500 | LCB 100 |
|---|---|---|---|---|
| Q8_0 | ๅทฒๅๅธๅนถๅฎๆ่ฟ็ซฏ้ช่ฏ | 164/198๏ผ82.83%๏ผ | 447/500๏ผ89.40%๏ผ | 74/100๏ผ74.00%๏ผ |
| Q6_K | ๅทฒๅๅธๅนถๅฎๆ่ฟ็ซฏ้ช่ฏ | 171/198๏ผ86.36%๏ผ | 440/500๏ผ88.00%๏ผ | 78/100๏ผ78.00%๏ผ |
| Q5-LynnStyle | ๅทฒๅๅธๅนถๅฎๆ่ฟ็ซฏ้ช่ฏ | 164/198๏ผ82.83%๏ผ | 438/500๏ผ87.60%๏ผ | 75/100๏ผ75.00%๏ผ |
| Q4-LynnStyle | ๅทฒๅๅธๅนถๅฎๆ่ฟ็ซฏ้ช่ฏ | 166/198๏ผ83.84%๏ผ | 443/500๏ผ88.60%๏ผ | 74/100๏ผ74.00%๏ผ |
| Q3-LynnStyle | ๅทฒๅๅธๅนถๅฎๆ่ฟ็ซฏ้ช่ฏ | 172/198๏ผ86.87%๏ผ | 435/500๏ผ87.00%๏ผ | 78/100๏ผ78.00%๏ผ |
| Q2-LynnStyle | ๅทฒๅๅธๅนถๅฎๆ่ฟ็ซฏ้ช่ฏ | 167/198๏ผ84.34%๏ผ | 416/500๏ผ83.20%๏ผ | 75/100๏ผ75.00%๏ผ |
Lynn Agent v0.87.0
Lynn Agent v0.87.0 ๅทฒ้็จๆฌ็ณปๅ Q2-LynnStyle / Q3-LynnStyle + DFlash2ใ่ฏฅ็ปๅๅทฒๅจ DGX Spark ๅฎๆต้่ฟ๏ผMac Apple Silicon/Intel ๅ
ฌ่ฏใWindows ๅฎ่ฃ
ๅ
่ฟ่กๆฃๆฅใไธคไป CIใไธไป main/tag ไธ่ดๆงใ23 ไธชๅ
ฌ็ฝๆไปถๅฎๆด SHA256 ไธ่ฟ็จ CLI ๅฎ่ฃ
ๅๅทฒ้่ฟใ
่ช็ถๆๆณ็ๆๅถๆๅฝฑไธๆๅ็ชๅ ๏ผ้ป่ฎคๅผๅฏ๏ผๆฌๅ้กถ้จโๆ ๅฝฑโๆฅ็ๅ ณ้ญ่ทฏๅพ๏ผ็นๅป็ด่พพ่ฎพ็ฝฎใๅๅฐๆๅ๏ผๅๅฐๅจๆๆๆๆถ้ๆญขใๅพ็ๅทฒๅนถๅ ฅๆไปถ็ญ้๏ผๆๆ ๆจกๆฟๅไปฃๅธธ้ฉปไปปๅกๆจกๅผ๏ผ็ฟป่ฏ็งปๅ ฅๆถๆฏ่ๅ๏ผไธๅฎถๅๆกๆนไธบๅฏ้ๆไปถ๏ผๅนถไฟฎๅคไผ่ฏ็ผ่พ็ฎๆ ไธๅๆญข้ขๅค็ใKimi Datasource ็ปง็ปญไฟ็ๅจ MCP ไธญ๏ผ็จๆท้่ช่กๆซ็ ็ปๅฝ่ชๅทฑ็่ดฆๅทใ
ๆฌ่ฝฎๅฎขๆท็ซฏๆดๆฐๆชๆนๅๆฌไปๆจกๅๆ้ใ้ๅๆไปถใๆต่ฏๅๆฐๆๆง่ฝๆๆ ใ
| ๅฎ่ฃ ๅ | ๅฝๅ ้ๅ | GitHub ๅค็จ |
|---|---|---|
| Mac Apple Silicon | ไธ่ฝฝ | ไธ่ฝฝ |
| Mac Intel | ไธ่ฝฝ | ไธ่ฝฝ |
| Windows | ไธ่ฝฝ | ไธ่ฝฝ |
ๅๅธ่ฎฐๅฝ๏ผGitHub ไธปไป ยท GitHub ๆงไป ยท Gitee ยท CLI ๅ
llama.cpp ็ๆฌไธๅฏๅจ่ฎพ็ฝฎ
ไฝฟ็จๅ
ๅซ 2026-08-27 ๅๅ
ฅ็ๅฎๆน DFlash2 ๆฏๆ ็ llama.cpp๏ผๆไบค b10f9ca58c89 ๆๅ
ผๅฎน็ๆฐ็ๆฌใไปฅไธไปๅฎๆนไปๆๅปบๆไฝๆฏๆ็ๆฌ๏ผ
git clone https://github.com/ggml-org/llama.cpp.git llama.cpp-efficientthink
git -C llama.cpp-efficientthink checkout b10f9ca58c89ccfc3653ac01e979dd085d582b76
cmake -S llama.cpp-efficientthink -B llama.cpp-efficientthink/build -DGGML_CUDA=ON
cmake --build llama.cpp-efficientthink/build --config Release -j 4 --target llama-server
export PATH="$PWD/llama.cpp-efficientthink/build/bin:$PATH"
llama-server --version
llama-server --help | grep -- draft-dflash
ไธๅๅๆกฃๅฝไปคๅจๆฌไปไธ่ฝฝๆ น็ฎๅฝๆง่ก๏ผไฝฟ็จไธปไป้ๅๆถๅ
cd GGUFใไธๆฌกๅช้ไธไธชๆกฃไฝใๅฏๅจไธไธชๆๅกใ็คบไพๆ็กฎ้็จ GPU 0๏ผๅชๆ่ช่ก้
็ฝฎๅคๅกๆถๆ็งป้คๆ่ฐๆด CUDA_VISIBLE_DEVICES=0ใ้ๅฎ่ฃ
ๆฏๆๆ็จ GPU ็ CUDA Toolkit ไธ็ผ่ฏๅจ๏ผGGUF ไธไพ่ต Spark FP8 ็ Docker ้ๅใ
ๅฎๆตๅธๅฑไธบ 8 slots ร 32,768 tokens๏ผๆปไธไธๆ 262144๏ผ--ctx-size 262144๏ผ๏ผไธๆฏๅ
ซๆงฝๅ
ฑ็จๆป่ฎก 32KใC4/C8 ่กจ็คบๅฏนๅ
ซๆงฝๆๅกๅๅบๅ/ๅ
ซไธชๅนถๅ่ฏทๆฑใๆพๅญ่พๅฐๆถๅบๅๆถ็ผฉๅฐๆงฝๆฐๅๆปไธไธๆ๏ผไพๅฆๅๆงฝ --parallel 1 --ctx-size 32768๏ผ่ฟ็ง้
็ฝฎไธ่ฝๆฒฟ็จๅกไธญ C4/C8 ็ๅๅ็ปๆใๆ้ๆไปถๅคงๅฐไธ็ญไบ่ฟ่กๆ้ๅ
จ้จๅ
ๅญใ
DFlash2 ๅฏ้ๆฉ็ฌ็ซ Q8_0 ๆ Q4_K_M draftใ--spec-type draft-dflashใ--spec-draft-n-max 7 ไธ --spec-draft-n-min 0๏ผไธๆฏๆจกๅๅ
็ฝฎ MTP ่ทฏ็บฟใ--jinja ๅๆพๅผ XH ๆจกๆฟๅๆฐไฟ็ๆ่ๆจกๅผใ้ป่ฎคๅช็ๅฌ 127.0.0.1:8080๏ผๅผๆพๅ
ถไปๆฅๅฃๅ่ฏท่ฎพ็ฝฎ่ฎฟ้ฎๆงๅถใ--mmproj ็จไบๅพๅ๏ผ็บฏๆๆฌๆๅกๅฏๅปๆ่ฏฅ่กใ
ๅฎๆนๅ่๏ผๆๅกๅๆฐใๆๆบ่งฃ็ ใCUDA ๆๅปบใไฟฎๆญฃๅ็ๅฝไปคๅฏนๅบๅฎๆนๅๆฐไธๅทฒๆตๆงฝไฝๅธๅฑ๏ผๆๅปบ่ฏดๆไธไปฃ่กจ้ๆฐ่ท่ฟๅ ญๆกฃ่ฝๅๆต่ฏใ
ไธบไปไน้ๆฉ DFlash2๏ผไปฅๅ้ๅ ๆไพ็ MTP ๆนๆก
ๆฌไปไปไผๅ ้็จ DFlash2๏ผไธป่ฆ็้ๅฎๆต่็จฟๆฅๅๅบฆๅ่พๅบ TPSใ ๅนถ่ก็ๆๅ้ๅใๆฏ่ฝฎ้ช่ฏๆฅๅๅคไธช token๏ผๆๅฉไบๅๅฐ้ token ้กบๅบ่งฃ็ ๅผ้ใๆนๆณ่ฏดๆ๏ผDFlash ไฝ่ ใ
ไพๅฆ๏ผQ8_0 ็ C4 ๅฎๆต่พพๅฐ 60.98% ๆฅๅ็ใ246 ่ๅ tok/s๏ผQ3-LynnStyle ็ C4 ่พพๅฐ 58.00%ใ269 ่ๅ tok/sใไธ้ขๆๅ้ๅๆกฃๅๅซๅๅบๆฐๆฎ๏ผC4 ไธบๅฎ็จๅนณ่กกๆกฃ๏ผC8 ไธบๆๅคงๅฎๆตๅๅๆกฃใๆฅๅๅบฆไธ TPS ๅๅณไบๆจกๅใ็กฌไปถใ่ฏทๆฑๅๅนถๅ๏ผ่ฟไบๆฐๆฎไธๆฏไธ MTP ็ๅๆกไปถๅฏน็ งใ
ๆฏๆกฃๅทฒ้ๅ ๆไพ Q4_0 ไธ Q8_0 MTP GGUF
Q2-LynnStyle ่ณ Q8_0 ็ๆฏไธช็ฎๅฝ้ฝๅ ๅซๅไธ็ป็ป่ฟ็ฌ็ซ้ช่ฏ็ MTP sidecar๏ผ
| ๆไปถ | ไฝ็จ | ๅคงๅฐ | SHA256 |
|---|---|---|---|
mtp-Qwen3.8-27B-Q4_0.gguf |
ๆดๅฐ็ Q4_0 MTP draft | 1,680,271,648 bytes | 051a1764cff8c4f3ee6ae8b00593a0364c7539c67fa50ffc58f3f96509fca38e |
mtp-Qwen3.8-27B-Q8_0.gguf |
ๆด้ซ็ฒพๅบฆ็ Q8_0 MTP draft | 3,164,006,688 bytes | cbf60a0c48b431bb61f1d49b8948dc88ac29c398d6dbdbbb2e6e89ef77eacc9a |
ไธคๆไปถๆฅ่ช ggml-org/Qwen3.8-27B-GGUF@0669b98607d47046c7c2b3f801011d54a08cfccf๏ผๅ่งฃๆไธบ 65 blocksใ18 tensors ็ Qwen MTP GGUFใDGX Spark ๅทฒๅฐไธค่
ๅๅซไธๆฌไป Q3-LynnStyle ไธปๆจกๅๅ็ๅฎๅ
ผๅฎนๆง smoke๏ผๅๆๅๅ ่ฝฝใHTTP 200ใ็ๆ้็ฉบๆ็ปๅ
ๅฎนๅนถๆญฃๅธธๅๆญขใ่ฏฅ็ปๆๅช่ฏๆๅ
ผๅฎนๆง๏ผไธๆฏๅป็ป็ MTP TPS ๆ่ฝๅๆ็ปฉใ
ๆฏๆฌกๅฏๅจๅช้ไธๆกๆๆบ่งฃ็ ่ทฏ็บฟ๏ผไธ่ฆๅๆถๅฏ็จ DFlash2 ๅ MTPใๆดๅฐ็ Q4_0 MTP ็คบไพ๏ผ
CUDA_VISIBLE_DEVICES=0 llama-server \
--model Q3-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q3-LynnStyle.gguf \
--model-draft Q3-LynnStyle/mtp-Qwen3.8-27B-Q4_0.gguf \
--spec-type draft-mtp \
--gpu-layers all --gpu-layers-draft all \
--parallel 1 --ctx-size 32768 --flash-attn on \
--jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
--host 127.0.0.1 --port 8080
ๅฆ้ Q8_0 MTP๏ผๅฐ draft ๆไปถๆฟๆขไธบ mtp-Qwen3.8-27B-Q8_0.ggufใๅชๆๅพๅๆๅกๆ้่ฆ้ขๅคๅ ๅ
ฅ่ฏฅๆกฃ็ --mmprojใ
ๅๆกฃ DFlash2 ๅนถๅๆง่ฝ
็ปไธๅฎๆตๅฃๅพ๏ผๅๅกไธปๆจกๅใDFlash2ใ--parallel 8 --ctx-size 262144๏ผ8 slots ร 32K๏ผ๏ผC1/C2/C4/C8๏ผ256-token ๅฎ้ฟๅๅ่พๅบ๏ผๆฏ่ฏทๆฑๆๅค 60 ็ง๏ผDFlash2 ไฝฟ็จ n_max=7ใn_min=0ใๅๅๅไธบ่ๅ completion tok/sใ
Q8_0
| ๅนถๅ | ่ๅๅๅ | DFlash2 ๆฅๅ็ | ๅนณๅๆฅๅ้ฟๅบฆ / 8 |
|---|---|---|---|
| C1 | 101.46 tok/s | 51.30% | 4.54 |
| C2 | 165.63 tok/s | 54.47% | 4.76 |
| C4 | 245.94 tok/s | 60.98% | 5.25 |
| C8 | 255.71 tok/s | 52.01% | 4.60 |
- ๅนถๅๅปบ่ฎฎ๏ผC4 ๅฎ็จๅนณ่กก๏ผC8 ๆๅคงๅๅใ
- ๆฏๆกฃ 15 ไธชๅฎ้ฟ่ฏทๆฑ๏ผ่ถ ๆถใ่ฏทๆฑ้่ฏฏใ็ฉบ่พๅบใๆ ็นๅๅกๅไธบ 0ใ
Q6_K
| ๅนถๅ | ่ๅๅๅ | DFlash2 ๆฅๅ็ | ๅนณๅๆฅๅ้ฟๅบฆ / 8 |
|---|---|---|---|
| C1 | 91.43 tok/s | 40.83% | 3.79 |
| C2 | 165.08 tok/s | 48.28% | 4.36 |
| C4 | 242.11 tok/s | 49.00% | 4.37 |
| C8 | 279.83 tok/s | 46.80% | 4.24 |
- ๅนถๅๅปบ่ฎฎ๏ผC4 ไฝๅปถ่ฟๅนณ่กก๏ผC8 ๆๅคงๅๅใ
- ๆด่ฝฎๆซๅไธญ๏ผ่ถ ๆถใ่ฏทๆฑ้่ฏฏใ็ฉบ่พๅบใๆ ็นๅๅกๅไธบ 0ใ
Q5-LynnStyle
| ๅนถๅ | ่ๅๅๅ | DFlash2 ๆฅๅ็ | ๅนณๅๆฅๅ้ฟๅบฆ / 8 |
|---|---|---|---|
| C1 | 94.64 tok/s | 40.87% | 3.81 |
| C2 | 158.71 tok/s | 39.98% | 3.78 |
| C4 | 238.06 tok/s | 48.22% | 4.35 |
| C8 | 275.73 tok/s | 48.38% | 4.34 |
- ๅนถๅๅปบ่ฎฎ๏ผC4 ๅฎ็จๅนณ่กก๏ผC8 ๆๅคงๅๅใ
- ๆด่ฝฎๆซๅไธญ๏ผ่ถ ๆถใ่ฏทๆฑ้่ฏฏใ็ฉบ่พๅบใๆ ็นๅๅกๅไธบ 0ใ
Q4-LynnStyle
| ๅนถๅ | ่ๅๅๅ | DFlash2 ๆฅๅ็ | ๅนณๅๆฅๅ้ฟๅบฆ / 8 |
|---|---|---|---|
| C1 | 97.96 tok/s | 38.83% | 3.70 |
| C2 | 149.27 tok/s | 36.15% | 3.51 |
| C4 | 256.52 tok/s | 47.69% | 4.30 |
| C8 | 263.50 tok/s | 42.88% | 3.95 |
- ๅนถๅๅปบ่ฎฎ๏ผC4 ๅฎ็จๅนณ่กก๏ผC8 ๆๅคงๅๅใ
- ๆด่ฝฎๆซๅไธญ๏ผ่ถ ๆถใ่ฏทๆฑ้่ฏฏใ็ฉบ่พๅบใๆ ็นๅๅกๅไธบ 0ใ
Q3-LynnStyle
| ๅนถๅ | ่ๅๅๅ | DFlash2 ๆฅๅ็ | ๅนณๅๆฅๅ้ฟๅบฆ / 8 |
|---|---|---|---|
| C1 | 112.97 tok/s | 47.22% | 4.25 |
| C2 | 210.78 tok/s | 57.22% | 5.00 |
| C4 | 269.32 tok/s | 58.00% | 5.02 |
| C8 | 286.16 tok/s | 52.41% | 4.63 |
- ๅนถๅๅปบ่ฎฎ๏ผC4 ไธบๅฎ็จๅนณ่กกๆกฃ๏ผC8 ไธบๆๅคงๅฎๆตๅๅใ
- ๆด่ฝฎๆซๅไธญ๏ผ่ถ ๆถใ่ฏทๆฑ้่ฏฏใ็ฉบ่พๅบใๆ ็นๅๅกๅไธบ 0ใ
่ฎญ็ปๆนๆณ
Qwen/Qwen3.8-27B โ ่ฝๅไฟๆ SFT โ ็ปๆญข่กไธบ SimPO โ ้ tensor FP32 delta ๅๅนถ โ BF16 โ GGUF ้ๅใ
- SFT๏ผ1,905 ๆกๆ ทๆฌ๏ผ1 epoch๏ผ239 optimizer steps๏ผeffective batch 8๏ผLoRA r=16ใalpha=32ใdropout=0.05๏ผLR 5e-6๏ผwarmup 12 steps๏ผseed 20260901ใ
- SimPO๏ผ110 ็ปๅๅฅฝๅฏนใ73 ไธชๅฏไธ prompt๏ผ5 optimizer steps๏ผbeta=1.0ใgamma=0.2ใpeak LR 5e-7๏ผLoRA r=16ใalpha=32ใdropout=0๏ผseed 20260903๏ผworld size 2๏ผFSDP full shardingใ
- ่ฎญ็ป็กฌไปถ๏ผๅ NVIDIA RTX PRO 6000 Blackwell Server Editionใ
ๆฌไปๅๅธๆๆฌไธป GGUFใ่ง่งๆๅฝฑไธ็ป่ฟ็ฌ็ซ้ช่ฏ็ DFlash2 draftใQ8_0 ไธ Q6_K ไฝฟ็จ Q8 ่ง่งๅก๏ผQ5-LynnStyleใQ4-LynnStyle ไธ Q3-LynnStyle ไฝฟ็จ Q6 ่ง่งๅกใๅ ญๆกฃ้ฝๆไพ Q8_0 ไธๆดๅฐ็ Q4_K_M DFlash2 draft๏ผๅนถๅฆๅซ Q4_0 ไธ Q8_0 MTP sidecarใๆฏๆฌกๅฏๅจๅช้ไธๆกๆๆบ่งฃ็ ่ทฏ็บฟ๏ผ่ฟไบ sidecar ้ฝไธๆฏไธปๆจกๅใ
ๅทฒๅๅธ GGUF ๆไปถ
Q8_0
ๆญฃๅผๅ จ้่ฏๆต๏ผGPQA 164/198๏ผ82.83%๏ผยท MMLU 447/500๏ผ89.40%๏ผยท LCB 74/100๏ผ74.00%๏ผใ
| ่ทฏๅพ | ไฝ็จ | ๅคงๅฐ |
|---|---|---|
Q8_0/Qwen3.8-27B-EfficientThink-SimPO-Q8_0.gguf |
Q8_0 ๆๆฌไธปๆจกๅ | 28,595,763,552 bytes |
Q8_0/mmproj-Qwen3.8-27B-Q8_0.gguf |
Q8 ่ง่งๆๅฝฑ | 629,247,008 bytes |
Q8_0/dflash2-qwen38-27b-Q8_0.gguf |
Q8_0 DFlash2 draft | 2,056,414,720 bytes |
Q8_0/dflash2-qwen38-27b-Q4_K_M.gguf |
ๅฏ้ Q4_K_M DFlash2 draft๏ผไธๆฏ MTP | 1,143,006,720 bytes |
Q8_0/mtp-Qwen3.8-27B-Q4_0.gguf |
ๅฏ้ Q4_0 MTP draft๏ผไธๆฏไธปๆจกๅๆ DFlash2 | 1,680,271,648 bytes |
Q8_0/mtp-Qwen3.8-27B-Q8_0.gguf |
ๅฏ้ Q8_0 MTP draft๏ผไธๆฏไธปๆจกๅๆ DFlash2 | 3,164,006,688 bytes |
Q8_0/SHA256SUMS |
ๆไปถๆ ก้ช | 396 bytes |
Q8_0/manifest.json |
็ปๆๅๆธ ๅ | 4,621 bytes |
ๅนถๅๅปบ่ฎฎ๏ผC4 ไธบๅฎ็จๅนณ่กกๆกฃ๏ผ246 tok/s๏ผ60.98% ๆฅๅ็๏ผ๏ผC8 ไธบๆๅคงๅฎๆตๅๅ๏ผ256 tok/s๏ผ52.01%๏ผใ
CUDA_VISIBLE_DEVICES=0 llama-server \
--model Q8_0/Qwen3.8-27B-EfficientThink-SimPO-Q8_0.gguf \
--mmproj Q8_0/mmproj-Qwen3.8-27B-Q8_0.gguf \
--model-draft Q8_0/dflash2-qwen38-27b-Q8_0.gguf \
--spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
--gpu-layers all --gpu-layers-draft all \
--parallel 8 --ctx-size 262144 --flash-attn on \
--jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
--host 127.0.0.1 --port 8080
Q5-LynnStyle
ๆญฃๅผๅ จ้่ฏๆต๏ผGPQA 164/198๏ผ82.83%๏ผยท MMLU 438/500๏ผ87.60%๏ผยท LCB 75/100๏ผ75.00%๏ผใLCB๏ผEasy 23/23ใMedium 29/31ใHard 23/46ใ
| ่ทฏๅพ | ไฝ็จ | ๅคงๅฐ |
|---|---|---|
Q5-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q5-LynnStyle.gguf |
Q5-LynnStyle ๆๆฌไธปๆจกๅ | 22,412,452,608 bytes |
Q5-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf |
Q6 ่ง่งๆๅฝฑ | 618,390,560 bytes |
Q5-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf |
Q8_0 DFlash2 draft | 2,056,414,720 bytes |
Q5-LynnStyle/dflash2-qwen38-27b-Q4_K_M.gguf |
ๅฏ้ Q4_K_M DFlash2 draft๏ผไธๆฏ MTP | 1,143,006,720 bytes |
Q5-LynnStyle/mtp-Qwen3.8-27B-Q4_0.gguf |
ๅฏ้ Q4_0 MTP draft๏ผไธๆฏไธปๆจกๅๆ DFlash2 | 1,680,271,648 bytes |
Q5-LynnStyle/mtp-Qwen3.8-27B-Q8_0.gguf |
ๅฏ้ Q8_0 MTP draft๏ผไธๆฏไธปๆจกๅๆ DFlash2 | 3,164,006,688 bytes |
Q5-LynnStyle/SHA256SUMS |
ๆไปถๆ ก้ช | 501 bytes |
Q5-LynnStyle/manifest.json |
็ปๆๅๆธ ๅ | 5,287 bytes |
ๅนถๅๅปบ่ฎฎ๏ผC4 ไธบๅฎ็จๅนณ่กกๆกฃ๏ผ238 tok/s๏ผ48.22% ๆฅๅ็๏ผ๏ผC8 ไธบๆๅคงๅฎๆตๅๅ๏ผ276 tok/s๏ผ48.38%๏ผใ
CUDA_VISIBLE_DEVICES=0 llama-server \
--model Q5-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q5-LynnStyle.gguf \
--mmproj Q5-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf \
--model-draft Q5-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf \
--spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
--gpu-layers all --gpu-layers-draft all \
--parallel 8 --ctx-size 262144 --flash-attn on \
--jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
--host 127.0.0.1 --port 8080
Q4-LynnStyle
ๆญฃๅผๅ จ้่ฏๆต๏ผGPQA 166/198๏ผ83.84%๏ผยท MMLU 443/500๏ผ88.60%๏ผยท LCB 74/100๏ผ74.00%๏ผใLCB๏ผEasy 23/23ใMedium 30/31ใHard 21/46๏ผ100 ไธชๅฏไธ้ข็ฎใerrors=0ใclean gate PASSใ
| ่ทฏๅพ | ไฝ็จ | ๅคงๅฐ |
|---|---|---|
Q4-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q4-LynnStyle.gguf |
Q4-LynnStyle ๆๆฌไธปๆจกๅ | 19,351,954,688 bytes |
Q4-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf |
Q6 ่ง่งๆๅฝฑ | 618,390,560 bytes |
Q4-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf |
ๅ ฑไบซ Q8_0 DFlash2 draft | 2,056,414,720 bytes |
Q4-LynnStyle/dflash2-qwen38-27b-Q4_K_M.gguf |
ๅฏ้ Q4_K_M DFlash2 draft๏ผไธๆฏ MTP | 1,143,006,720 bytes |
Q4-LynnStyle/mtp-Qwen3.8-27B-Q4_0.gguf |
ๅฏ้ Q4_0 MTP draft๏ผไธๆฏไธปๆจกๅๆ DFlash2 | 1,680,271,648 bytes |
Q4-LynnStyle/mtp-Qwen3.8-27B-Q8_0.gguf |
ๅฏ้ Q8_0 MTP draft๏ผไธๆฏไธปๆจกๅๆ DFlash2 | 3,164,006,688 bytes |
Q4-LynnStyle/SHA256SUMS |
ๆไปถๆ ก้ช | 501 bytes |
Q4-LynnStyle/manifest.json |
็ปๆๅๆธ ๅ | 5,439 bytes |
| ๅนถๅ | ่ๅๅๅ | DFlash2 ๆฅๅ็ | ๅนณๅๆฅๅ้ฟๅบฆ / 8 |
|---|---|---|---|
| C1 | 97.96 tok/s | 38.83% | 3.70 |
| C2 | 149.27 tok/s | 36.15% | 3.51 |
| C4 | 256.52 tok/s | 47.69% | 4.30 |
| C8 | 263.50 tok/s | 42.88% | 3.95 |
ๅนถๅๅปบ่ฎฎ๏ผC4 ไธบๅฎ็จๅนณ่กกๆกฃ๏ผ257 tok/s๏ผ47.69% ๆฅๅ็๏ผ๏ผC8 ไธบๆๅคงๅฎๆตๅๅ๏ผ264 tok/s๏ผ42.88%๏ผใๆด่ฝฎๆซๅไธญ๏ผ่ถ ๆถใ่ฏทๆฑ้่ฏฏใ็ฉบ่พๅบใๆ ็นๅๅกๅไธบ 0ใ
CUDA_VISIBLE_DEVICES=0 llama-server \
--model Q4-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q4-LynnStyle.gguf \
--mmproj Q4-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf \
--model-draft Q4-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf \
--spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
--gpu-layers all --gpu-layers-draft all \
--parallel 8 --ctx-size 262144 --flash-attn on \
--jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
--host 127.0.0.1 --port 8080
Q3-LynnStyle
ๆญฃๅผๅ จ้่ฏๆต๏ผGPQA 172/198๏ผ86.87%๏ผยท MMLU 435/500๏ผ87.00%๏ผยท LCB 78/100๏ผ78.00%๏ผใGPQA ไธบ 198 ไธชๅฏไธ้ข็ฎใerrors=0ใstrict loop=0๏ผMMLU errors=0ใparse fail=0๏ผLCB๏ผEasy 23/23ใMedium 28/31ใHard 27/46๏ผ100 ไธชๅฏไธ้ข็ฎใerrors=0ใๆๆๆช้่ฟๆ ทๆฌๅไฟ็ๅจๅๆฏใ
| ่ทฏๅพ | ไฝ็จ | ๅคงๅฐ |
|---|---|---|
Q3-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q3-LynnStyle.gguf |
Q3-LynnStyle ๆๆฌไธปๆจกๅ | 17,032,184,960 bytes |
Q3-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf |
Q6 ่ง่งๆๅฝฑ | 618,390,560 bytes |
Q3-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf |
ๅ ฑไบซ Q8_0 DFlash2 draft | 2,056,414,720 bytes |
Q3-LynnStyle/dflash2-qwen38-27b-Q4_K_M.gguf |
ๅฏ้ Q4_K_M DFlash2 draft๏ผไธๆฏ MTP | 1,143,006,720 bytes |
Q3-LynnStyle/mtp-Qwen3.8-27B-Q4_0.gguf |
ๅฏ้ Q4_0 MTP draft๏ผไธๆฏไธปๆจกๅๆ DFlash2 | 1,680,271,648 bytes |
Q3-LynnStyle/mtp-Qwen3.8-27B-Q8_0.gguf |
ๅฏ้ Q8_0 MTP draft๏ผไธๆฏไธปๆจกๅๆ DFlash2 | 3,164,006,688 bytes |
Q3-LynnStyle/SHA256SUMS |
ๆไปถๆ ก้ช | 501 bytes |
Q3-LynnStyle/manifest.json |
็ปๆๅๆธ ๅ | 5,526 bytes |
ๅนถๅๅปบ่ฎฎ๏ผC4 ไธบๅฎ็จๅนณ่กกๆกฃ๏ผ269 tok/s๏ผ58.00% ๆฅๅ็๏ผ๏ผC8 ไธบๆๅคงๅฎๆตๅๅ๏ผ286 tok/s๏ผ52.41%๏ผใๆด่ฝฎๆซๅไธญ๏ผ่ถ ๆถใ่ฏทๆฑ้่ฏฏใ็ฉบ่พๅบใๆ ็นๅๅกๅไธบ 0ใ
CUDA_VISIBLE_DEVICES=0 llama-server \
--model Q3-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q3-LynnStyle.gguf \
--mmproj Q3-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf \
--model-draft Q3-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf \
--spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
--gpu-layers all --gpu-layers-draft all \
--parallel 8 --ctx-size 262144 --flash-attn on \
--jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
--host 127.0.0.1 --port 8080
Q2-LynnStyle
ๆญฃๅผๅ จ้ๆ็ปฉ๏ผGPQA 167/198๏ผ84.34%๏ผยท MMLU 416/500๏ผ83.20%๏ผยท LCB 75/100๏ผ75.00%๏ผใLCB๏ผEasy 23/23ใMedium 29/31ใHard 23/46๏ผ23 ไธช 32K ้ฟๅบฆ็ปๆๅไปฅๅคฑ่ดฅ่ฎกๅ ฅใ
LCB ๆ็ป่งๅพ็ฑๅๅง่ฟ่กไธญไฟๆ้ๅญ่ไธๅ็ 99 ้ข๏ผๅ ไธ Lynn ๆๆๅฏน atcoder:arc181_b:81 ็ไธๆฌก็ฒพ็กฎ่กฅๆต็ปๆใ่ฏฅ้ขๅ่ฟ่กๅ็ๆตๅผ JSON ้่ฏฏ๏ผ่กฅๆตไฝฟ็จๅไธๆจกๅใrunnerใ่ฏทๆฑไธๅๆฐ๏ผๅฎๆดๆถๅฐ DONE๏ผไฝไปๅจ 32K ็ปๆไธๆฒกๆๆไบคไปฃ็ ๏ผๅ ๆญคๆปๅไปไธบ 75/100ใ่ฏฅ็ปๆไธๆฏไธๆฌก่ฟ็ปญๆ ๆ
้่ฟ่กใ
| ่ทฏๅพ | ไฝ็จ | ๅคงๅฐ |
|---|---|---|
Q2-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q2-LynnStyle.gguf |
Q2-LynnStyle ๆๆฌไธปๆจกๅ | 12,999,977,600 bytes |
Q2-LynnStyle/mmproj-Qwen3.8-27B-Q4_K_M.gguf |
Q4_K_M ่ง่งๆๅฝฑ | 522,293,024 bytes |
Q2-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf |
ๅฏ้ Q8_0 DFlash2 draft๏ผไธๆฏ MTP | 2,056,414,720 bytes |
Q2-LynnStyle/dflash2-qwen38-27b-Q4_K_M.gguf |
ๅฏ้ Q4_K_M DFlash2 draft๏ผไธๆฏ MTP | 1,143,006,720 bytes |
Q2-LynnStyle/mtp-Qwen3.8-27B-Q4_0.gguf |
ๅฏ้ Q4_0 MTP draft๏ผไธๆฏไธปๆจกๅๆ DFlash2 | 1,680,271,648 bytes |
Q2-LynnStyle/mtp-Qwen3.8-27B-Q8_0.gguf |
ๅฏ้ Q8_0 MTP draft๏ผไธๆฏไธปๆจกๅๆ DFlash2 | 3,164,006,688 bytes |
Q2-LynnStyle/SHA256SUMS |
ๆไปถๆ ก้ช | 406 bytes |
Q2-LynnStyle/manifest.json |
็ปๆๅๆธ ๅ | 5,504 bytes |
Q2-LynnStyle ไฝฟ็จ GSQ-RCO ๆททๅ็ฒพๅบฆ้ๅ๏ผๅนถไปฅ IQ ๆฐๅผ็ปๅไผๅ 70 ไธช้ๅฎ MLP ็ฉ้ต๏ผattentionใhead ไธ SSM tensor ๆชๅๆฐๅผ็ปๅใ่ฅๅ ๅญๅ ่ฎธ๏ผQ3-LynnStyle ๅๆด้ซๆกฃไฝไปๆฏๆด็จณๅฆฅ็่ดจ้้ๆฉใๆญค็ฒพ็กฎ 12,999,977,600-byte ๆๅปบๆฒกๆๅป็ป TPS๏ผไธ่ฝๅฅ็จๅทฒๆทๆฑฐๅทฅ็จ Q2 ็ๅๅใ
CUDA_VISIBLE_DEVICES=0 llama-server \
--model Q2-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q2-LynnStyle.gguf \
--mmproj Q2-LynnStyle/mmproj-Qwen3.8-27B-Q4_K_M.gguf \
--model-draft Q2-LynnStyle/dflash2-qwen38-27b-Q4_K_M.gguf \
--spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
--gpu-layers all --gpu-layers-draft all \
--parallel 4 --ctx-size 131072 --flash-attn on \
--jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
--host 127.0.0.1 --port 8080
Q8_0ใQ6_KใQ5-LynnStyleใQ4-LynnStyleใQ3-LynnStyle ไธ Q2-LynnStyle ็ไธป GGUF ๅไธบ GGUF v3ใqwen35ใ64 ๅฑใ851 tensors๏ผไธปๆไปถไธๅ
ๅต่ง่งๆ MTP tensor๏ผๅๅทฒ้่ฟ่ฟ็ซฏๅญ่/SHA256 ๅ่ฏปใๅฝๅๅๅธ็ Q8_0 DFlash2 ๅทฒ้่ฟ้ๆ่งฃๆไธ Q2-LynnStyle ไปฃ่กจๆง่ฟ่ก smoke๏ผๅ
ถไฝๅๆกฃๅฎๆดๅนถๅ้ๆตๅพ
่กฅ๏ผๆง DFlash2 ๆฐๆฎไธๅฝๅ ไบ่ฏฅ Q8_0 ๆไปถใ
LynnStyle ๆททๅ็ฒพๅบฆ้ๅ
Q5/Q4/Q3 LynnStyle ๅไปๆ็ป SimPO F16 ๆๆฌไธปๅนฒ็ดๆฅ้ๅ๏ผไฝฟ็จๅป็ป็ 512-chunk imatrix๏ผ496 entries๏ผ๏ผไธๅไบๆฌก requantizeใไฟๆค็ญ็ฅ็ปผๅ T10 REAPใ้ฆๅฐพๅฑ/SSM ไฟๆคไธ full-attentionใSSMใFFN ็้ tensor ๅฎก่ฎก๏ผ่งๅๅฒ็ชๆถ้็จๆด้ซ็ฒพๅบฆใ
- Q5-LynnStyle๏ผๅทฒๅๅธ๏ผ๏ผๅบ็ก Q5๏ผL0โ3ใL60โ63 ไธๅ
ณ้ฎ REAP ๅฑ 50/52/54/56/58 ไฝฟ็จ Q8๏ผ้ซๆๆๅฑ 21/42/44/46/48/49/53/57/59 ไฝฟ็จ Q6๏ผfull-attention q/k/v/o ไธ้ๅฎไธญๅฑ
ssm_outไฝฟ็จ Q6ใ498 ไธช่ฝฌๆข tensor๏ผQ8_0 104ใQ6_K 186ใQ5_K 208ใ - Q4-LynnStyle๏ผๅทฒๅๅธ๏ผ๏ผๅบ็ก Q4๏ผL0โ3ใL60โ63 ไฝฟ็จ Q8๏ผๅ
ณ้ฎ่ฉ้จๅฑ 50/52/54/56โ59 ไฝฟ็จ Q6๏ผ้ซๆๆๅฑ 21/42/44/46/48/49/53 ไฝฟ็จ Q5๏ผfull-attention q/k/v/o ไธ้ๅฎ
ssm_outไฝฟ็จ Q6ใๅๅธ๏ผQ8 62ใQ6 156ใQ5 172ใQ4 108ใ - Q3-LynnStyle๏ผๅทฒๅๅธ๏ผ๏ผๅบ็ก Q3๏ผL0โ3ใL60โ63 ไฝฟ็จ Q6๏ผL4โ6ใ50/52/54/56โ59 ไฝฟ็จ Q5๏ผ้ซๆๆๅฑ 21/42/44/46/48/49/53 ไฝฟ็จ Q4๏ผfull-attention ไธไธญๅฑ
ssm_outไฝฟ็จ Q5ใๅๅธ๏ผQ8 2ใQ6 62ใQ5 157ใQ4 199ใQ3 78ใ
Hub ไพง่พนๆ ๅฏ่ฝไป
ๆพ็คบ GGUF ๅไน general.file_type๏ผๆ ๆณๅฎๆด่กจ่พพ่ชๅฎไน้ tensor ๆททๅ็ฒพๅบฆ๏ผไปฅๅ็ฎๅฝ manifest.jsonใSHA256SUMS ไธๅฎ้
GGUF parse ไธบๅใ
Q6_K
ๆญฃๅผๅ จ้่ฏๆต๏ผGPQA 171/198๏ผ86.36%๏ผยท MMLU 440/500๏ผ88.00%๏ผยท LCB 78/100๏ผ78.00%๏ผใ
LCB ้พๅบฆๅๅธ๏ผEasy 23/23ใMedium 30/31ใHard 25/46ใ
| ๅนถๅ | ่ๅๅๅ | DFlash2 ๆฅๅ็ | ๅนณๅๆฅๅ้ฟๅบฆ / 8 | ๆๅคง่ฏทๆฑๆถ้ฟ |
|---|---|---|---|---|
| C1 | 91 tok/s | 40.83% | 3.79 | 2.80s |
| C2 | 165 tok/s | 48.28% | 4.36 | 3.10s |
| C4๏ผไฝๅปถ่ฟๅนณ่กกๆกฃ๏ผ | 242 tok/s | 49.00% | 4.37 | 4.23s |
| C8๏ผๆๅคงๅๅ๏ผ | 280 tok/s | 46.80% | 4.24 | 7.31s |
ๅๆกฃๅบๅฎ่พๅบๆต่ฏ็่ฏทๆฑ้่ฏฏใ่ถ ๆถใ็ฉบ่พๅบๅๆ ็นๅๅกๅไธบ 0ใๅๅๆต่ฏไธไปฃๆฟๆญฃๅผ่ดจ้่ฏๆตใ
ๆไปถ
| ่ทฏๅพ | ไฝ็จ | ๅคงๅฐ |
|---|---|---|
Q6_K/Qwen3.8-27B-EfficientThink-SimPO-Q6_K.gguf |
Q6_K ๆๆฌไธปๆจกๅ | 22,082,530,048 bytes |
Q6_K/mmproj-Qwen3.8-27B-Q8_0.gguf |
Q8 ่ง่งๆๅฝฑ | 629,247,008 bytes |
Q6_K/dflash2-qwen38-27b-Q8_0.gguf |
Q8_0 DFlash2 draft | 2,056,414,720 bytes |
Q6_K/dflash2-qwen38-27b-Q4_K_M.gguf |
ๅฏ้ Q4_K_M DFlash2 draft๏ผไธๆฏ MTP | 1,143,006,720 bytes |
Q6_K/mtp-Qwen3.8-27B-Q4_0.gguf |
ๅฏ้ Q4_0 MTP draft๏ผไธๆฏไธปๆจกๅๆ DFlash2 | 1,680,271,648 bytes |
Q6_K/mtp-Qwen3.8-27B-Q8_0.gguf |
ๅฏ้ Q8_0 MTP draft๏ผไธๆฏไธปๆจกๅๆ DFlash2 | 3,164,006,688 bytes |
ๆจ่ๅฏๅจ
CUDA_VISIBLE_DEVICES=0 llama-server \
--model Q6_K/Qwen3.8-27B-EfficientThink-SimPO-Q6_K.gguf \
--mmproj Q6_K/mmproj-Qwen3.8-27B-Q8_0.gguf \
--model-draft Q6_K/dflash2-qwen38-27b-Q8_0.gguf \
--spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
--gpu-layers all --gpu-layers-draft all \
--parallel 8 --ctx-size 262144 --flash-attn on \
--jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
--host 127.0.0.1 --port 8080
- Downloads last month
- 87,129

