Qwen3.8-27B EfficientThink GGUF

Qwen3.8-27B EfficientThink capability and reasoning comparison

No strict loops were observed in the reviewed Q2โ€“Q8 evaluations.

BF16 / FP8 main repository: nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2

Measured quantization-tier overview

Capability scores are shown only after a formal full-suite freeze, with every non-passing sample retained in the denominator. Q8_0, Q6_K, Q5-LynnStyle, Q4-LynnStyle, Q3-LynnStyle, and Q2-LynnStyle are published and have passed four-repository remote path, byte-size, and LFS-SHA256 readback.

Tier Repository state GPQA 198 MMLU 500 LCB 100
Q8_0 Published and remotely verified 164/198 (82.83%) 447/500 (89.40%) 74/100 (74.00%)
Q6_K Published and remotely verified 171/198 (86.36%) 440/500 (88.00%) 78/100 (78.00%)
Q5-LynnStyle Published and remotely verified 164/198 (82.83%) 438/500 (87.60%) 75/100 (75.00%)
Q4-LynnStyle Published and remotely verified 166/198 (83.84%) 443/500 (88.60%) 74/100 (74.00%)
Q3-LynnStyle Published and remotely verified 172/198 (86.87%) 435/500 (87.00%) 78/100 (78.00%)
Q2-LynnStyle Published and remotely verified 167/198 (84.34%) 416/500 (83.20%) 75/100 (75.00%)

Lynn Agent v0.87.0

Lynn Agent v0.87.0 uses this release's Q2-LynnStyle / Q3-LynnStyle + DFlash2 packages. The pairing passed runtime validation on DGX Spark; notarized Mac Apple Silicon and Intel builds, the Windows installer runtime check, CI in both repositories, matching main heads and tags across all three release repositories, complete SHA256 verification for 23 public files, and a remote CLI installation all passed.

Natural moving tree shadows and soft window light, enabled by default. Hover over Shadows for the off switch location, or click to open Settings. Playback pauses in the background and stays still with reduced motion. Images now participate in file filtering; slash templates replace persistent task mode; translation moved into the message menu; Expert Roundtable is now an optional plugin; and session edit targeting and stop preprocessing were fixed. Kimi Datasource remains available under MCP and requires users to scan the QR code and sign in with their own account.

This client update did not change any model weight, quantized artifact, benchmark score, or performance metric in this repository.

Installer China mirror GitHub fallback
Mac Apple Silicon Download Download
Mac Intel Download Download
Windows Download Download

Release records: primary GitHub repository ยท legacy GitHub repository ยท Gitee ยท CLI package

llama.cpp prerequisites and launch settings

Use an official llama.cpp build containing DFlash2 support, merged 2026-08-27, commit b10f9ca58c89 or a compatible newer build. A reproducible minimum build from the official repository is:

git clone https://github.com/ggml-org/llama.cpp.git llama.cpp-efficientthink
git -C llama.cpp-efficientthink checkout b10f9ca58c89ccfc3653ac01e979dd085d582b76
cmake -S llama.cpp-efficientthink -B llama.cpp-efficientthink/build -DGGML_CUDA=ON
cmake --build llama.cpp-efficientthink/build --config Release -j 4 --target llama-server
export PATH="$PWD/llama.cpp-efficientthink/build/bin:$PATH"
llama-server --version
llama-server --help | grep -- draft-dflash

Run the tier commands below from this repository's download root. If using the main repository's mirror, first cd GGUF. Choose one tier/server at a time. Each command explicitly selects GPU 0; remove or adjust CUDA_VISIBLE_DEVICES=0 only for your own multi-GPU configuration. CUDA Toolkit and a compiler supported by the chosen GPU are prerequisites; the Spark FP8 Docker image is not required for GGUF.

The measured layout is 8 slots ร— 32,768 tokens = --ctx-size 262144 total, not 32K total shared by eight slots. C4/C8 means four/eight active client requests against the eight-slot server. On smaller-memory hardware reduce both slots and total context (for example one slot with --parallel 1 --ctx-size 32768); those settings do not inherit the published C4/C8 throughput. Weight size alone is not the total runtime memory requirement.

DFlash2 can use either the separate Q8_0 or Q4_K_M draft, --spec-type draft-dflash, --spec-draft-n-max 7, and --spec-draft-n-min 0. It is not the in-model MTP path. --jinja and explicit XH template arguments retain the model's thinking mode. The examples bind to 127.0.0.1:8080; expose a different interface only with appropriate access controls. The --mmproj sidecar is for images; omit that line for text-only serving.

Official references: server options, speculative decoding, CUDA build. These corrected commands match the documented flags and measured slot layout; the build instructions are not a new six-tier benchmark run.

Why DFlash2, and bundled MTP alternatives

DFlash2 remains our preferred speculative decoder for its measured draft acceptance and output TPS. Parallel block drafting and acceptance of multiple tokens per verification help reduce sequential decoding overhead. Method reference: DFlash authors.

For example, the measured Q8_0 C4 configuration achieved 60.98% acceptance and 246 aggregate tok/s; Q3-LynnStyle C4 achieved 58.00% and 269 aggregate tok/s. The tables below report each tier independently: C4 is the practical balance and C8 the maximum measured throughput. Acceptance and TPS depend on the model, hardware, workload, and concurrency; these results are not a matched MTP comparison.

Q4_0 and Q8_0 MTP GGUF files are bundled

Every Q2-LynnStyle through Q8_0 directory now contains the same two independently verified MTP sidecars:

File Role Size SHA256
mtp-Qwen3.8-27B-Q4_0.gguf Smaller Q4_0 MTP draft 1,680,271,648 bytes 051a1764cff8c4f3ee6ae8b00593a0364c7539c67fa50ffc58f3f96509fca38e
mtp-Qwen3.8-27B-Q8_0.gguf Higher-precision Q8_0 MTP draft 3,164,006,688 bytes cbf60a0c48b431bb61f1d49b8948dc88ac29c398d6dbdbbb2e6e89ef77eacc9a

Both are derived from ggml-org/Qwen3.8-27B-GGUF@0669b98607d47046c7c2b3f801011d54a08cfccf and parse as Qwen MTP GGUFs with 65 blocks and 18 tensors. A real DGX Spark smoke paired each file with the published Q3-LynnStyle main model: both loaded, returned HTTP 200, generated non-empty final content, and stopped normally. This is a compatibility smoke, not a frozen MTP TPS or capability benchmark.

Choose one speculative path per launch. Do not combine DFlash2 and MTP. Example with the smaller MTP draft:

CUDA_VISIBLE_DEVICES=0 llama-server \
  --model Q3-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q3-LynnStyle.gguf \
  --model-draft Q3-LynnStyle/mtp-Qwen3.8-27B-Q4_0.gguf \
  --spec-type draft-mtp \
  --gpu-layers all --gpu-layers-draft all \
  --parallel 1 --ctx-size 32768 --flash-attn on \
  --jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
  --host 127.0.0.1 --port 8080

Substitute mtp-Qwen3.8-27B-Q8_0.gguf for the Q8_0 MTP option. Add the tier's --mmproj only when serving images.

Per-tier DFlash2 concurrency performance

Common measured protocol: one GPU for the main model, DFlash2, --parallel 8 --ctx-size 262144 (8 slots ร— 32K), C1/C2/C4/C8, fixed 256-token pressure outputs, and a 60-second request ceiling. DFlash2 used n_max=7 and n_min=0. Throughput is aggregate completion tok/s.

Q8_0

Concurrency Aggregate throughput DFlash2 acceptance Mean accepted length / 8
C1 101.46 tok/s 51.30% 4.54
C2 165.63 tok/s 54.47% 4.76
C4 245.94 tok/s 60.98% 5.25
C8 255.71 tok/s 52.01% 4.60
  • Concurrency guidance: C4 practical balance; C8 maximum throughput.
  • 15 fixed-length requests per tier; zero timeout, request error, empty output, or punctuation collapse.

Q6_K

Concurrency Aggregate throughput DFlash2 acceptance Mean accepted length / 8
C1 91.43 tok/s 40.83% 3.79
C2 165.08 tok/s 48.28% 4.36
C4 242.11 tok/s 49.00% 4.37
C8 279.83 tok/s 46.80% 4.24
  • Concurrency guidance: C4 lower-latency balance; C8 maximum throughput.
  • Zero timeout, request error, empty output, or punctuation collapse across the sweep.

Q5-LynnStyle

Concurrency Aggregate throughput DFlash2 acceptance Mean accepted length / 8
C1 94.64 tok/s 40.87% 3.81
C2 158.71 tok/s 39.98% 3.78
C4 238.06 tok/s 48.22% 4.35
C8 275.73 tok/s 48.38% 4.34
  • Concurrency guidance: C4 practical balance; C8 maximum throughput.
  • Zero timeout, request error, empty output, or punctuation collapse across the sweep.

Q4-LynnStyle

Concurrency Aggregate throughput DFlash2 acceptance Mean accepted length / 8
C1 97.96 tok/s 38.83% 3.70
C2 149.27 tok/s 36.15% 3.51
C4 256.52 tok/s 47.69% 4.30
C8 263.50 tok/s 42.88% 3.95
  • Concurrency guidance: C4 practical balance; C8 maximum throughput.
  • Zero timeout, request error, empty output, or punctuation collapse across the sweep.

Q3-LynnStyle

Concurrency Aggregate throughput DFlash2 acceptance Mean accepted length / 8
C1 112.97 tok/s 47.22% 4.25
C2 210.78 tok/s 57.22% 5.00
C4 269.32 tok/s 58.00% 5.02
C8 286.16 tok/s 52.41% 4.63
  • Concurrency guidance: C4 is the practical balance; C8 is the maximum measured throughput.
  • The full sweep had zero timeout, request error, empty output, or punctuation collapse.

Training method

Qwen/Qwen3.8-27B โ†’ capability-preserving SFT โ†’ terminal-behavior SimPO โ†’ per-tensor FP32 delta merge โ†’ BF16 โ†’ GGUF quantization.

  • SFT: 1,905 samples, 1 epoch, 239 optimizer steps, effective batch 8; LoRA r=16, alpha=32, dropout=0.05; LR 5e-6, 12 warmup steps, seed 20260901.
  • SimPO: 110 preference pairs over 73 unique prompts, 5 optimizer steps; beta=1.0, gamma=0.2, peak LR 5e-7; LoRA r=16, alpha=32, dropout=0, seed 20260903, world size 2, FSDP full sharding.
  • Training hardware: 2x NVIDIA RTX PRO 6000 Blackwell Server Edition.

This repository publishes text-main GGUF files with independently verified vision and DFlash2 sidecars. Q8_0 and Q6_K use the Q8 vision projector; Q5-LynnStyle, Q4-LynnStyle, and Q3-LynnStyle use the Q6 vision projector. All six tiers provide Q8_0 and smaller Q4_K_M DFlash2 drafts plus Q4_0 and Q8_0 MTP sidecars. Choose exactly one speculative path per launch; none of these sidecars is a main model.

Published GGUF files

Q8_0

Formal full-suite evaluation: GPQA 164/198 (82.83%) ยท MMLU 447/500 (89.40%) ยท LCB 74/100 (74.00%).

Path Role Size
Q8_0/Qwen3.8-27B-EfficientThink-SimPO-Q8_0.gguf Q8_0 text-main model 28,595,763,552 bytes
Q8_0/mmproj-Qwen3.8-27B-Q8_0.gguf Q8 vision projector 629,247,008 bytes
Q8_0/dflash2-qwen38-27b-Q8_0.gguf Q8_0 DFlash2 draft 2,056,414,720 bytes
Q8_0/dflash2-qwen38-27b-Q4_K_M.gguf Optional Q4_K_M DFlash2 draft; not MTP 1,143,006,720 bytes
Q8_0/mtp-Qwen3.8-27B-Q4_0.gguf Optional Q4_0 MTP draft; not a main model or DFlash2 1,680,271,648 bytes
Q8_0/mtp-Qwen3.8-27B-Q8_0.gguf Optional Q8_0 MTP draft; not a main model or DFlash2 3,164,006,688 bytes
Q8_0/SHA256SUMS Checksums 396 bytes
Q8_0/manifest.json Structured manifest 4,621 bytes

Concurrency guidance: C4 is the practical balance (246 tok/s, 60.98% acceptance); C8 is the maximum measured throughput (256 tok/s, 52.01%).

CUDA_VISIBLE_DEVICES=0 llama-server \
  --model Q8_0/Qwen3.8-27B-EfficientThink-SimPO-Q8_0.gguf \
  --mmproj Q8_0/mmproj-Qwen3.8-27B-Q8_0.gguf \
  --model-draft Q8_0/dflash2-qwen38-27b-Q8_0.gguf \
  --spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
  --gpu-layers all --gpu-layers-draft all \
  --parallel 8 --ctx-size 262144 --flash-attn on \
  --jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
  --host 127.0.0.1 --port 8080

Q5-LynnStyle

Formal full-suite evaluation: GPQA 164/198 (82.83%) ยท MMLU 438/500 (87.60%) ยท LCB 75/100 (75.00%). LCB: Easy 23/23, Medium 29/31, Hard 23/46.

Path Role Size
Q5-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q5-LynnStyle.gguf Q5-LynnStyle text-main model 22,412,452,608 bytes
Q5-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf Q6 vision projector 618,390,560 bytes
Q5-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf Q8_0 DFlash2 draft 2,056,414,720 bytes
Q5-LynnStyle/dflash2-qwen38-27b-Q4_K_M.gguf Optional Q4_K_M DFlash2 draft; not MTP 1,143,006,720 bytes
Q5-LynnStyle/mtp-Qwen3.8-27B-Q4_0.gguf Optional Q4_0 MTP draft; not a main model or DFlash2 1,680,271,648 bytes
Q5-LynnStyle/mtp-Qwen3.8-27B-Q8_0.gguf Optional Q8_0 MTP draft; not a main model or DFlash2 3,164,006,688 bytes
Q5-LynnStyle/SHA256SUMS Checksums 501 bytes
Q5-LynnStyle/manifest.json Structured manifest 5,287 bytes

Concurrency guidance: C4 is the practical balance (238 tok/s, 48.22% acceptance); C8 is the maximum measured throughput (276 tok/s, 48.38%).

CUDA_VISIBLE_DEVICES=0 llama-server \
  --model Q5-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q5-LynnStyle.gguf \
  --mmproj Q5-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf \
  --model-draft Q5-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf \
  --spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
  --gpu-layers all --gpu-layers-draft all \
  --parallel 8 --ctx-size 262144 --flash-attn on \
  --jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
  --host 127.0.0.1 --port 8080

Q4-LynnStyle

Formal full-suite evaluation: GPQA 166/198 (83.84%) ยท MMLU 443/500 (88.60%) ยท LCB 74/100 (74.00%). LCB: Easy 23/23, Medium 30/31, Hard 21/46; 100 unique tasks, errors=0, clean gate PASS.

Path Role Size
Q4-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q4-LynnStyle.gguf Q4-LynnStyle text-main model 19,351,954,688 bytes
Q4-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf Q6 vision projector 618,390,560 bytes
Q4-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf Shared Q8_0 DFlash2 draft 2,056,414,720 bytes
Q4-LynnStyle/dflash2-qwen38-27b-Q4_K_M.gguf Optional Q4_K_M DFlash2 draft; not MTP 1,143,006,720 bytes
Q4-LynnStyle/mtp-Qwen3.8-27B-Q4_0.gguf Optional Q4_0 MTP draft; not a main model or DFlash2 1,680,271,648 bytes
Q4-LynnStyle/mtp-Qwen3.8-27B-Q8_0.gguf Optional Q8_0 MTP draft; not a main model or DFlash2 3,164,006,688 bytes
Q4-LynnStyle/SHA256SUMS Checksums 501 bytes
Q4-LynnStyle/manifest.json Structured manifest 5,439 bytes
Concurrency Aggregate throughput DFlash2 acceptance Mean accepted length / 8
C1 97.96 tok/s 38.83% 3.70
C2 149.27 tok/s 36.15% 3.51
C4 256.52 tok/s 47.69% 4.30
C8 263.50 tok/s 42.88% 3.95

Concurrency guidance: C4 is the practical balance (257 tok/s, 47.69% acceptance); C8 is the maximum measured throughput (264 tok/s, 42.88%). The full sweep had zero timeout, request error, empty output, or punctuation collapse.

CUDA_VISIBLE_DEVICES=0 llama-server \
  --model Q4-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q4-LynnStyle.gguf \
  --mmproj Q4-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf \
  --model-draft Q4-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf \
  --spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
  --gpu-layers all --gpu-layers-draft all \
  --parallel 8 --ctx-size 262144 --flash-attn on \
  --jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
  --host 127.0.0.1 --port 8080

Q3-LynnStyle

Formal full-suite evaluation: GPQA 172/198 (86.87%) ยท MMLU 435/500 (87.00%) ยท LCB 78/100 (78.00%). GPQA used 198 unique items with errors=0 and strict loop=0; MMLU had errors=0 and parse fail=0; LCB: Easy 23/23, Medium 28/31, Hard 27/46, with 100 unique tasks and errors=0. Every non-passing sample remains in the denominator.

Path Role Size
Q3-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q3-LynnStyle.gguf Q3-LynnStyle text-main model 17,032,184,960 bytes
Q3-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf Q6 vision projector 618,390,560 bytes
Q3-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf Shared Q8_0 DFlash2 draft 2,056,414,720 bytes
Q3-LynnStyle/dflash2-qwen38-27b-Q4_K_M.gguf Optional Q4_K_M DFlash2 draft; not MTP 1,143,006,720 bytes
Q3-LynnStyle/mtp-Qwen3.8-27B-Q4_0.gguf Optional Q4_0 MTP draft; not a main model or DFlash2 1,680,271,648 bytes
Q3-LynnStyle/mtp-Qwen3.8-27B-Q8_0.gguf Optional Q8_0 MTP draft; not a main model or DFlash2 3,164,006,688 bytes
Q3-LynnStyle/SHA256SUMS Checksums 501 bytes
Q3-LynnStyle/manifest.json Structured manifest 5,526 bytes

Concurrency guidance: C4 is the practical balance (269 tok/s, 58.00% acceptance); C8 is the maximum measured throughput (286 tok/s, 52.41%). The full sweep had zero timeout, request error, empty output, or punctuation collapse.

CUDA_VISIBLE_DEVICES=0 llama-server \
  --model Q3-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q3-LynnStyle.gguf \
  --mmproj Q3-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf \
  --model-draft Q3-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf \
  --spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
  --gpu-layers all --gpu-layers-draft all \
  --parallel 8 --ctx-size 262144 --flash-attn on \
  --jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
  --host 127.0.0.1 --port 8080

Q2-LynnStyle

Formal full-suite scores: GPQA 167/198 (84.34%) ยท MMLU 416/500 (83.20%) ยท LCB 75/100 (75.00%). LCB: Easy 23/23, Medium 29/31, Hard 23/46; all 23 length-limited 32K outputs remain failures in the denominator.

The final LCB view combines the byte-identical 99 unaffected rows from the original run with one Lynn-authorized exact retry of atcoder:arc181_b:81. The original request ended in a streaming JSON error. The retry used the same model, runner, request, and parameters and received a complete DONE, but again reached 32K with no submitted code, so the score remains 75/100. This is not represented as one uninterrupted error-free run.

Path Role Size
Q2-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q2-LynnStyle.gguf Q2-LynnStyle text-main model 12,999,977,600 bytes
Q2-LynnStyle/mmproj-Qwen3.8-27B-Q4_K_M.gguf Q4_K_M vision projector 522,293,024 bytes
Q2-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf Optional Q8_0 DFlash2 draft; not MTP 2,056,414,720 bytes
Q2-LynnStyle/dflash2-qwen38-27b-Q4_K_M.gguf Optional Q4_K_M DFlash2 draft; not MTP 1,143,006,720 bytes
Q2-LynnStyle/mtp-Qwen3.8-27B-Q4_0.gguf Optional Q4_0 MTP draft; not a main model or DFlash2 1,680,271,648 bytes
Q2-LynnStyle/mtp-Qwen3.8-27B-Q8_0.gguf Optional Q8_0 MTP draft; not a main model or DFlash2 3,164,006,688 bytes
Q2-LynnStyle/SHA256SUMS Checksums 406 bytes
Q2-LynnStyle/manifest.json Structured manifest 5,504 bytes

Q2-LynnStyle uses GSQ-RCO mixed-precision quantization with IQ numerical refinement over 70 selected MLP matrices; attention, head, and SSM tensors were not numerically refined. If memory permits, Q3-LynnStyle or a higher tier remains the safer quality choice. This exact 12,999,977,600-byte build has no frozen TPS result, so throughput from the retired engineering Q2 is not reused.

CUDA_VISIBLE_DEVICES=0 llama-server \
  --model Q2-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q2-LynnStyle.gguf \
  --mmproj Q2-LynnStyle/mmproj-Qwen3.8-27B-Q4_K_M.gguf \
  --model-draft Q2-LynnStyle/dflash2-qwen38-27b-Q4_K_M.gguf \
  --spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
  --gpu-layers all --gpu-layers-draft all \
  --parallel 4 --ctx-size 131072 --flash-attn on \
  --jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
  --host 127.0.0.1 --port 8080

The Q8_0, Q6_K, Q5-LynnStyle, Q4-LynnStyle, Q3-LynnStyle, and Q2-LynnStyle main GGUF files are GGUF v3, qwen35, 64-layer, 851-tensor text trunks with no embedded vision or MTP tensors; all passed remote byte/SHA256 readback. The published Q8_0 DFlash2 passed static parsing and a representative Q2-LynnStyle runtime smoke. Full per-tier concurrency retests remain pending, and historical DFlash2 grids are not attributed to this Q8_0 file.

LynnStyle mixed-precision quantization

Q5/Q4/Q3 LynnStyle tiers are quantized directly from the final SimPO F16 text trunk with the frozen 512-chunk imatrix (496 entries), without requantization. Protection combines T10 REAP, first/last-layer and SSM guards, plus per-tensor full-attention, SSM, and FFN audits; higher precision wins when rules overlap.

  • Q5-LynnStyle (published): Q5 base; Q8 for L0โ€“3, L60โ€“63, and critical REAP layers 50/52/54/56/58; Q6 for high-sensitivity layers 21/42/44/46/48/49/53/57/59, full-attention q/k/v/o, and selected middle ssm_out. Of 498 converted tensors: Q8_0 104, Q6_K 186, Q5_K 208.
  • Q4-LynnStyle (published): Q4 base; Q8 for L0โ€“3 and L60โ€“63; Q6 for critical shoulder layers 50/52/54/56โ€“59; Q5 for high layers 21/42/44/46/48/49/53; Q6 for full-attention q/k/v/o and selected ssm_out. Distribution: Q8 62, Q6 156, Q5 172, Q4 108.
  • Q3-LynnStyle (published): Q3 base; Q6 for L0โ€“3 and L60โ€“63; Q5 for L4โ€“6 and 50/52/54/56โ€“59; Q4 for high layers 21/42/44/46/48/49/53; Q5 for full-attention and middle ssm_out. Distribution: Q8 2, Q6 62, Q5 157, Q4 199, Q3 78.

Hub sidebars may expose only nominal GGUF general.file_type, which cannot fully describe custom per-tensor mixed precision. Treat each tier's manifest.json, SHA256SUMS, and actual GGUF parse as authoritative.

Q6_K

Formal full-suite evaluation: GPQA 171/198 (86.36%) ยท MMLU 440/500 (88.00%) ยท LCB 78/100 (78.00%).

LCB difficulty split: Easy 23/23, Medium 30/31, Hard 25/46.

Concurrency Aggregate throughput DFlash2 acceptance Mean accepted length / 8 Maximum request time
C1 91 tok/s 40.83% 3.79 2.80s
C2 165 tok/s 48.28% 4.36 3.10s
C4 (lower-latency balance) 242 tok/s 49.00% 4.37 4.23s
C8 (maximum throughput) 280 tok/s 46.80% 4.24 7.31s

Request errors, timeouts, empty outputs, and punctuation collapse were zero at all four tested concurrency levels. This serving sweep characterizes throughput and speculative acceptance; it does not replace formal quality evaluation.

Files

Path Role Size
Q6_K/Qwen3.8-27B-EfficientThink-SimPO-Q6_K.gguf Q6_K text-main model 22,082,530,048 bytes
Q6_K/mmproj-Qwen3.8-27B-Q8_0.gguf Q8 vision projector 629,247,008 bytes
Q6_K/dflash2-qwen38-27b-Q8_0.gguf Q8_0 DFlash2 draft 2,056,414,720 bytes
Q6_K/dflash2-qwen38-27b-Q4_K_M.gguf Optional Q4_K_M DFlash2 draft; not MTP 1,143,006,720 bytes
Q6_K/mtp-Qwen3.8-27B-Q4_0.gguf Optional Q4_0 MTP draft; not a main model or DFlash2 1,680,271,648 bytes
Q6_K/mtp-Qwen3.8-27B-Q8_0.gguf Optional Q8_0 MTP draft; not a main model or DFlash2 3,164,006,688 bytes

Recommended launch

CUDA_VISIBLE_DEVICES=0 llama-server \
  --model Q6_K/Qwen3.8-27B-EfficientThink-SimPO-Q6_K.gguf \
  --mmproj Q6_K/mmproj-Qwen3.8-27B-Q8_0.gguf \
  --model-draft Q6_K/dflash2-qwen38-27b-Q8_0.gguf \
  --spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
  --gpu-layers all --gpu-layers-draft all \
  --parallel 8 --ctx-size 262144 --flash-attn on \
  --jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
  --host 127.0.0.1 --port 8080

ไธญๆ–‡่ฏดๆ˜Ž

Qwen3.8-27B EfficientThink ่ƒฝๅŠ›ไธŽๆ€่€ƒๅฏนๆฏ”

ๅทฒๅฎกๆŸฅ็š„ Q2โ€“Q8 ๆต‹่ฏ„ไธญๆœชๅ‘็Žฐไธฅๆ ผๆญปๅพช็Žฏใ€‚

BF16 / FP8 ไธปไป“๏ผšnerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2

้‡ๅŒ–ๆกฃไฝๅฎžๆต‹ๆ€ป่งˆ

่ƒฝๅŠ›ๅˆ†ๆ•ฐๅชๅˆ—ๆญฃๅผๅ…จ้‡ๅ†ป็ป“็ป“ๆžœ๏ผŒๆ‰€ๆœ‰ๆœช้€š่ฟ‡ๆ ทๆœฌๅ‡ไฟ็•™ๅœจๅˆ†ๆฏใ€‚Q8_0ใ€Q6_Kใ€Q5-LynnStyleใ€Q4-LynnStyleใ€Q3-LynnStyle ไธŽ Q2-LynnStyle ๅ‡ๅทฒๅ‘ๅธƒ๏ผŒๅนถๅฎŒๆˆๅ››ไป“่ฟœ็ซฏ่ทฏๅพ„ใ€ๅญ—่Š‚ๅ’Œ LFS SHA256 ๅ›ž่ฏปใ€‚

ๆกฃไฝ ไป“ๅบ“็Šถๆ€ GPQA 198 MMLU 500 LCB 100
Q8_0 ๅทฒๅ‘ๅธƒๅนถๅฎŒๆˆ่ฟœ็ซฏ้ชŒ่ฏ 164/198๏ผˆ82.83%๏ผ‰ 447/500๏ผˆ89.40%๏ผ‰ 74/100๏ผˆ74.00%๏ผ‰
Q6_K ๅทฒๅ‘ๅธƒๅนถๅฎŒๆˆ่ฟœ็ซฏ้ชŒ่ฏ 171/198๏ผˆ86.36%๏ผ‰ 440/500๏ผˆ88.00%๏ผ‰ 78/100๏ผˆ78.00%๏ผ‰
Q5-LynnStyle ๅทฒๅ‘ๅธƒๅนถๅฎŒๆˆ่ฟœ็ซฏ้ชŒ่ฏ 164/198๏ผˆ82.83%๏ผ‰ 438/500๏ผˆ87.60%๏ผ‰ 75/100๏ผˆ75.00%๏ผ‰
Q4-LynnStyle ๅทฒๅ‘ๅธƒๅนถๅฎŒๆˆ่ฟœ็ซฏ้ชŒ่ฏ 166/198๏ผˆ83.84%๏ผ‰ 443/500๏ผˆ88.60%๏ผ‰ 74/100๏ผˆ74.00%๏ผ‰
Q3-LynnStyle ๅทฒๅ‘ๅธƒๅนถๅฎŒๆˆ่ฟœ็ซฏ้ชŒ่ฏ 172/198๏ผˆ86.87%๏ผ‰ 435/500๏ผˆ87.00%๏ผ‰ 78/100๏ผˆ78.00%๏ผ‰
Q2-LynnStyle ๅทฒๅ‘ๅธƒๅนถๅฎŒๆˆ่ฟœ็ซฏ้ชŒ่ฏ 167/198๏ผˆ84.34%๏ผ‰ 416/500๏ผˆ83.20%๏ผ‰ 75/100๏ผˆ75.00%๏ผ‰

Lynn Agent v0.87.0

Lynn Agent v0.87.0 ๅทฒ้‡‡็”จๆœฌ็ณปๅˆ— Q2-LynnStyle / Q3-LynnStyle + DFlash2ใ€‚่ฏฅ็ป„ๅˆๅทฒๅœจ DGX Spark ๅฎžๆต‹้€š่ฟ‡๏ผ›Mac Apple Silicon/Intel ๅ…ฌ่ฏใ€Windows ๅฎ‰่ฃ…ๅŒ…่ฟ่กŒๆฃ€ๆŸฅใ€ไธคไป“ CIใ€ไธ‰ไป“ main/tag ไธ€่‡ดๆ€งใ€23 ไธชๅ…ฌ็ฝ‘ๆ–‡ไปถๅฎŒๆ•ด SHA256 ไธŽ่ฟœ็จ‹ CLI ๅฎ‰่ฃ…ๅ‡ๅทฒ้€š่ฟ‡ใ€‚

่‡ช็„ถๆ‘‡ๆ›ณ็š„ๆžๅถๆŠ•ๅฝฑไธŽๆŸ”ๅ’Œ็ช—ๅ…‰๏ผŒ้ป˜่ฎคๅผ€ๅฏ๏ผ›ๆ‚ฌๅœ้กถ้ƒจโ€˜ๆ ‘ๅฝฑโ€™ๆŸฅ็œ‹ๅ…ณ้—ญ่ทฏๅพ„๏ผŒ็‚นๅ‡ป็›ด่พพ่ฎพ็ฝฎใ€‚ๅŽๅฐๆš‚ๅœ๏ผŒๅ‡ๅฐ‘ๅŠจๆ€ๆ•ˆๆžœๆ—ถ้™ๆญขใ€‚ๅ›พ็‰‡ๅทฒๅนถๅ…ฅๆ–‡ไปถ็ญ›้€‰๏ผŒๆ–œๆ ๆจกๆฟๅ–ไปฃๅธธ้ฉปไปปๅŠกๆจกๅผ๏ผŒ็ฟป่ฏ‘็งปๅ…ฅๆถˆๆฏ่œๅ•๏ผŒไธ“ๅฎถๅœ†ๆกŒๆ”นไธบๅฏ้€‰ๆ’ไปถ๏ผŒๅนถไฟฎๅคไผš่ฏ็ผ–่พ‘็›ฎๆ ‡ไธŽๅœๆญข้ข„ๅค„็†ใ€‚Kimi Datasource ็ปง็ปญไฟ็•™ๅœจ MCP ไธญ๏ผŒ็”จๆˆท้œ€่‡ช่กŒๆ‰ซ็ ็™ปๅฝ•่‡ชๅทฑ็š„่ดฆๅทใ€‚

ๆœฌ่ฝฎๅฎขๆˆท็ซฏๆ›ดๆ–ฐๆœชๆ”นๅ˜ๆœฌไป“ๆจกๅž‹ๆƒ้‡ใ€้‡ๅŒ–ๆ–‡ไปถใ€ๆต‹่ฏ„ๅˆ†ๆ•ฐๆˆ–ๆ€ง่ƒฝๆŒ‡ๆ ‡ใ€‚

ๅฎ‰่ฃ…ๅŒ… ๅ›ฝๅ†…้•œๅƒ GitHub ๅค‡็”จ
Mac Apple Silicon ไธ‹่ฝฝ ไธ‹่ฝฝ
Mac Intel ไธ‹่ฝฝ ไธ‹่ฝฝ
Windows ไธ‹่ฝฝ ไธ‹่ฝฝ

ๅ‘ๅธƒ่ฎฐๅฝ•๏ผšGitHub ไธปไป“ ยท GitHub ๆ—งไป“ ยท Gitee ยท CLI ๅŒ…

llama.cpp ็‰ˆๆœฌไธŽๅฏๅŠจ่ฎพ็ฝฎ

ไฝฟ็”จๅŒ…ๅซ 2026-08-27 ๅˆๅ…ฅ็š„ๅฎ˜ๆ–น DFlash2 ๆ”ฏๆŒ ็š„ llama.cpp๏ผšๆไบค b10f9ca58c89 ๆˆ–ๅ…ผๅฎน็š„ๆ–ฐ็‰ˆๆœฌใ€‚ไปฅไธ‹ไปŽๅฎ˜ๆ–นไป“ๆž„ๅปบๆœ€ไฝŽๆ”ฏๆŒ็‰ˆๆœฌ๏ผš

git clone https://github.com/ggml-org/llama.cpp.git llama.cpp-efficientthink
git -C llama.cpp-efficientthink checkout b10f9ca58c89ccfc3653ac01e979dd085d582b76
cmake -S llama.cpp-efficientthink -B llama.cpp-efficientthink/build -DGGML_CUDA=ON
cmake --build llama.cpp-efficientthink/build --config Release -j 4 --target llama-server
export PATH="$PWD/llama.cpp-efficientthink/build/bin:$PATH"
llama-server --version
llama-server --help | grep -- draft-dflash

ไธ‹ๅˆ—ๅˆ†ๆกฃๅ‘ฝไปคๅœจๆœฌไป“ไธ‹่ฝฝๆ น็›ฎๅฝ•ๆ‰ง่กŒ๏ผ›ไฝฟ็”จไธปไป“้•œๅƒๆ—ถๅ…ˆ cd GGUFใ€‚ไธ€ๆฌกๅช้€‰ไธ€ไธชๆกฃไฝใ€ๅฏๅŠจไธ€ไธชๆœๅŠกใ€‚็คบไพ‹ๆ˜Ž็กฎ้€‰็”จ GPU 0๏ผ›ๅชๆœ‰่‡ช่กŒ้…็ฝฎๅคšๅกๆ—ถๆ‰็งป้™คๆˆ–่ฐƒๆ•ด CUDA_VISIBLE_DEVICES=0ใ€‚้œ€ๅฎ‰่ฃ…ๆ”ฏๆŒๆ‰€็”จ GPU ็š„ CUDA Toolkit ไธŽ็ผ–่ฏ‘ๅ™จ๏ผ›GGUF ไธไพ่ต– Spark FP8 ็š„ Docker ้•œๅƒใ€‚

ๅฎžๆต‹ๅธƒๅฑ€ไธบ 8 slots ร— 32,768 tokens๏ผŒๆ€ปไธŠไธ‹ๆ–‡ 262144๏ผˆ--ctx-size 262144๏ผ‰๏ผŒไธๆ˜ฏๅ…ซๆงฝๅ…ฑ็”จๆ€ป่ฎก 32Kใ€‚C4/C8 ่กจ็คบๅฏนๅ…ซๆงฝๆœๅŠกๅ‘ๅ‡บๅ››/ๅ…ซไธชๅนถๅ‘่ฏทๆฑ‚ใ€‚ๆ˜พๅญ˜่พƒๅฐๆ—ถๅบ”ๅŒๆ—ถ็ผฉๅฐๆงฝๆ•ฐๅ’Œๆ€ปไธŠไธ‹ๆ–‡๏ผŒไพ‹ๅฆ‚ๅ•ๆงฝ --parallel 1 --ctx-size 32768๏ผ›่ฟ™็ง้…็ฝฎไธ่ƒฝๆฒฟ็”จๅกไธญ C4/C8 ็š„ๅžๅ็ป“ๆžœใ€‚ๆƒ้‡ๆ–‡ไปถๅคงๅฐไธ็ญ‰ไบŽ่ฟ่กŒๆ‰€้œ€ๅ…จ้ƒจๅ†…ๅญ˜ใ€‚

DFlash2 ๅฏ้€‰ๆ‹ฉ็‹ฌ็ซ‹ Q8_0 ๆˆ– Q4_K_M draftใ€--spec-type draft-dflashใ€--spec-draft-n-max 7 ไธŽ --spec-draft-n-min 0๏ผŒไธๆ˜ฏๆจกๅž‹ๅ†…็ฝฎ MTP ่ทฏ็บฟใ€‚--jinja ๅ’Œๆ˜พๅผ XH ๆจกๆฟๅ‚ๆ•ฐไฟ็•™ๆ€่€ƒๆจกๅผใ€‚้ป˜่ฎคๅช็›‘ๅฌ 127.0.0.1:8080๏ผŒๅผ€ๆ”พๅ…ถไป–ๆŽฅๅฃๅ‰่ฏท่ฎพ็ฝฎ่ฎฟ้—ฎๆŽงๅˆถใ€‚--mmproj ็”จไบŽๅ›พๅƒ๏ผ›็บฏๆ–‡ๆœฌๆœๅŠกๅฏๅŽปๆމ่ฏฅ่กŒใ€‚

ๅฎ˜ๆ–นๅ‚่€ƒ๏ผšๆœๅŠกๅ‚ๆ•ฐใ€ๆŠ•ๆœบ่งฃ็ ใ€CUDA ๆž„ๅปบใ€‚ไฟฎๆญฃๅŽ็š„ๅ‘ฝไปคๅฏนๅบ”ๅฎ˜ๆ–นๅ‚ๆ•ฐไธŽๅทฒๆต‹ๆงฝไฝๅธƒๅฑ€๏ผ›ๆž„ๅปบ่ฏดๆ˜Žไธไปฃ่กจ้‡ๆ–ฐ่ท‘่ฟ‡ๅ…ญๆกฃ่ƒฝๅŠ›ๆต‹่ฏ„ใ€‚

ไธบไป€ไนˆ้€‰ๆ‹ฉ DFlash2๏ผŒไปฅๅŠ้šๅŒ…ๆไพ›็š„ MTP ๆ–นๆกˆ

ๆœฌไป“ไปไผ˜ๅ…ˆ้‡‡็”จ DFlash2๏ผŒไธป่ฆ็œ‹้‡ๅฎžๆต‹่‰็จฟๆŽฅๅ—ๅบฆๅ’Œ่พ“ๅ‡บ TPSใ€‚ ๅนถ่กŒ็”Ÿๆˆๅ€™้€‰ๅ—ใ€ๆฏ่ฝฎ้ชŒ่ฏๆŽฅๅ—ๅคšไธช token๏ผŒๆœ‰ๅŠฉไบŽๅ‡ๅฐ‘้€ token ้กบๅบ่งฃ็ ๅผ€้”€ใ€‚ๆ–นๆณ•่ฏดๆ˜Ž๏ผšDFlash ไฝœ่€…ใ€‚

ไพ‹ๅฆ‚๏ผŒQ8_0 ็š„ C4 ๅฎžๆต‹่พพๅˆฐ 60.98% ๆŽฅๅ—็އใ€246 ่šๅˆ tok/s๏ผ›Q3-LynnStyle ็š„ C4 ่พพๅˆฐ 58.00%ใ€269 ่šๅˆ tok/sใ€‚ไธ‹้ขๆŒ‰ๅ„้‡ๅŒ–ๆกฃๅˆ†ๅˆซๅˆ—ๅ‡บๆ•ฐๆฎ๏ผšC4 ไธบๅฎž็”จๅนณ่กกๆกฃ๏ผŒC8 ไธบๆœ€ๅคงๅฎžๆต‹ๅžๅๆกฃใ€‚ๆŽฅๅ—ๅบฆไธŽ TPS ๅ–ๅ†ณไบŽๆจกๅž‹ใ€็กฌไปถใ€่ฏทๆฑ‚ๅ’Œๅนถๅ‘๏ผ›่ฟ™ไบ›ๆ•ฐๆฎไธๆ˜ฏไธŽ MTP ็š„ๅŒๆกไปถๅฏน็…งใ€‚

ๆฏๆกฃๅทฒ้šๅŒ…ๆไพ› Q4_0 ไธŽ Q8_0 MTP GGUF

Q2-LynnStyle ่‡ณ Q8_0 ็š„ๆฏไธช็›ฎๅฝ•้ƒฝๅŒ…ๅซๅŒไธ€็ป„็ป่ฟ‡็‹ฌ็ซ‹้ชŒ่ฏ็š„ MTP sidecar๏ผš

ๆ–‡ไปถ ไฝœ็”จ ๅคงๅฐ SHA256
mtp-Qwen3.8-27B-Q4_0.gguf ๆ›ดๅฐ็š„ Q4_0 MTP draft 1,680,271,648 bytes 051a1764cff8c4f3ee6ae8b00593a0364c7539c67fa50ffc58f3f96509fca38e
mtp-Qwen3.8-27B-Q8_0.gguf ๆ›ด้ซ˜็ฒพๅบฆ็š„ Q8_0 MTP draft 3,164,006,688 bytes cbf60a0c48b431bb61f1d49b8948dc88ac29c398d6dbdbbb2e6e89ef77eacc9a

ไธคๆ–‡ไปถๆฅ่‡ช ggml-org/Qwen3.8-27B-GGUF@0669b98607d47046c7c2b3f801011d54a08cfccf๏ผŒๅ‡่งฃๆžไธบ 65 blocksใ€18 tensors ็š„ Qwen MTP GGUFใ€‚DGX Spark ๅทฒๅฐ†ไธค่€…ๅˆ†ๅˆซไธŽๆœฌไป“ Q3-LynnStyle ไธปๆจกๅž‹ๅš็œŸๅฎžๅ…ผๅฎนๆ€ง smoke๏ผšๅ‡ๆˆๅŠŸๅŠ ่ฝฝใ€HTTP 200ใ€็”Ÿๆˆ้ž็ฉบๆœ€็ปˆๅ†…ๅฎนๅนถๆญฃๅธธๅœๆญขใ€‚่ฏฅ็ป“ๆžœๅช่ฏๆ˜Žๅ…ผๅฎนๆ€ง๏ผŒไธๆ˜ฏๅ†ป็ป“็š„ MTP TPS ๆˆ–่ƒฝๅŠ›ๆˆ็ปฉใ€‚

ๆฏๆฌกๅฏๅŠจๅช้€‰ไธ€ๆกๆŠ•ๆœบ่งฃ็ ่ทฏ็บฟ๏ผŒไธ่ฆๅŒๆ—ถๅฏ็”จ DFlash2 ๅ’Œ MTPใ€‚ๆ›ดๅฐ็š„ Q4_0 MTP ็คบไพ‹๏ผš

CUDA_VISIBLE_DEVICES=0 llama-server \
  --model Q3-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q3-LynnStyle.gguf \
  --model-draft Q3-LynnStyle/mtp-Qwen3.8-27B-Q4_0.gguf \
  --spec-type draft-mtp \
  --gpu-layers all --gpu-layers-draft all \
  --parallel 1 --ctx-size 32768 --flash-attn on \
  --jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
  --host 127.0.0.1 --port 8080

ๅฆ‚้œ€ Q8_0 MTP๏ผŒๅฐ† draft ๆ–‡ไปถๆ›ฟๆขไธบ mtp-Qwen3.8-27B-Q8_0.ggufใ€‚ๅชๆœ‰ๅ›พๅƒๆœๅŠกๆ‰้œ€่ฆ้ขๅค–ๅŠ ๅ…ฅ่ฏฅๆกฃ็š„ --mmprojใ€‚

ๅˆ†ๆกฃ DFlash2 ๅนถๅ‘ๆ€ง่ƒฝ

็ปŸไธ€ๅฎžๆต‹ๅฃๅพ„๏ผšๅ•ๅกไธปๆจกๅž‹ใ€DFlash2ใ€--parallel 8 --ctx-size 262144๏ผˆ8 slots ร— 32K๏ผ‰๏ผŒC1/C2/C4/C8๏ผŒ256-token ๅฎš้•ฟๅŽ‹ๅŠ›่พ“ๅ‡บ๏ผŒๆฏ่ฏทๆฑ‚ๆœ€ๅคš 60 ็ง’๏ผ›DFlash2 ไฝฟ็”จ n_max=7ใ€n_min=0ใ€‚ๅžๅๅ‡ไธบ่šๅˆ completion tok/sใ€‚

Q8_0

ๅนถๅ‘ ่šๅˆๅžๅ DFlash2 ๆŽฅๅ—็އ ๅนณๅ‡ๆŽฅๅ—้•ฟๅบฆ / 8
C1 101.46 tok/s 51.30% 4.54
C2 165.63 tok/s 54.47% 4.76
C4 245.94 tok/s 60.98% 5.25
C8 255.71 tok/s 52.01% 4.60
  • ๅนถๅ‘ๅปบ่ฎฎ๏ผšC4 ๅฎž็”จๅนณ่กก๏ผ›C8 ๆœ€ๅคงๅžๅใ€‚
  • ๆฏๆกฃ 15 ไธชๅฎš้•ฟ่ฏทๆฑ‚๏ผ›่ถ…ๆ—ถใ€่ฏทๆฑ‚้”™่ฏฏใ€็ฉบ่พ“ๅ‡บใ€ๆ ‡็‚นๅๅกŒๅ‡ไธบ 0ใ€‚

Q6_K

ๅนถๅ‘ ่šๅˆๅžๅ DFlash2 ๆŽฅๅ—็އ ๅนณๅ‡ๆŽฅๅ—้•ฟๅบฆ / 8
C1 91.43 tok/s 40.83% 3.79
C2 165.08 tok/s 48.28% 4.36
C4 242.11 tok/s 49.00% 4.37
C8 279.83 tok/s 46.80% 4.24
  • ๅนถๅ‘ๅปบ่ฎฎ๏ผšC4 ไฝŽๅปถ่ฟŸๅนณ่กก๏ผ›C8 ๆœ€ๅคงๅžๅใ€‚
  • ๆ•ด่ฝฎๆ‰ซๅ‚ไธญ๏ผŒ่ถ…ๆ—ถใ€่ฏทๆฑ‚้”™่ฏฏใ€็ฉบ่พ“ๅ‡บใ€ๆ ‡็‚นๅๅกŒๅ‡ไธบ 0ใ€‚

Q5-LynnStyle

ๅนถๅ‘ ่šๅˆๅžๅ DFlash2 ๆŽฅๅ—็އ ๅนณๅ‡ๆŽฅๅ—้•ฟๅบฆ / 8
C1 94.64 tok/s 40.87% 3.81
C2 158.71 tok/s 39.98% 3.78
C4 238.06 tok/s 48.22% 4.35
C8 275.73 tok/s 48.38% 4.34
  • ๅนถๅ‘ๅปบ่ฎฎ๏ผšC4 ๅฎž็”จๅนณ่กก๏ผ›C8 ๆœ€ๅคงๅžๅใ€‚
  • ๆ•ด่ฝฎๆ‰ซๅ‚ไธญ๏ผŒ่ถ…ๆ—ถใ€่ฏทๆฑ‚้”™่ฏฏใ€็ฉบ่พ“ๅ‡บใ€ๆ ‡็‚นๅๅกŒๅ‡ไธบ 0ใ€‚

Q4-LynnStyle

ๅนถๅ‘ ่šๅˆๅžๅ DFlash2 ๆŽฅๅ—็އ ๅนณๅ‡ๆŽฅๅ—้•ฟๅบฆ / 8
C1 97.96 tok/s 38.83% 3.70
C2 149.27 tok/s 36.15% 3.51
C4 256.52 tok/s 47.69% 4.30
C8 263.50 tok/s 42.88% 3.95
  • ๅนถๅ‘ๅปบ่ฎฎ๏ผšC4 ๅฎž็”จๅนณ่กก๏ผ›C8 ๆœ€ๅคงๅžๅใ€‚
  • ๆ•ด่ฝฎๆ‰ซๅ‚ไธญ๏ผŒ่ถ…ๆ—ถใ€่ฏทๆฑ‚้”™่ฏฏใ€็ฉบ่พ“ๅ‡บใ€ๆ ‡็‚นๅๅกŒๅ‡ไธบ 0ใ€‚

Q3-LynnStyle

ๅนถๅ‘ ่šๅˆๅžๅ DFlash2 ๆŽฅๅ—็އ ๅนณๅ‡ๆŽฅๅ—้•ฟๅบฆ / 8
C1 112.97 tok/s 47.22% 4.25
C2 210.78 tok/s 57.22% 5.00
C4 269.32 tok/s 58.00% 5.02
C8 286.16 tok/s 52.41% 4.63
  • ๅนถๅ‘ๅปบ่ฎฎ๏ผšC4 ไธบๅฎž็”จๅนณ่กกๆกฃ๏ผ›C8 ไธบๆœ€ๅคงๅฎžๆต‹ๅžๅใ€‚
  • ๆ•ด่ฝฎๆ‰ซๅ‚ไธญ๏ผŒ่ถ…ๆ—ถใ€่ฏทๆฑ‚้”™่ฏฏใ€็ฉบ่พ“ๅ‡บใ€ๆ ‡็‚นๅๅกŒๅ‡ไธบ 0ใ€‚

่ฎญ็ปƒๆ–นๆณ•

Qwen/Qwen3.8-27B โ†’ ่ƒฝๅŠ›ไฟๆŒ SFT โ†’ ็ปˆๆญข่กŒไธบ SimPO โ†’ ้€ tensor FP32 delta ๅˆๅนถ โ†’ BF16 โ†’ GGUF ้‡ๅŒ–ใ€‚

  • SFT๏ผš1,905 ๆกๆ ทๆœฌ๏ผŒ1 epoch๏ผŒ239 optimizer steps๏ผŒeffective batch 8๏ผ›LoRA r=16ใ€alpha=32ใ€dropout=0.05๏ผ›LR 5e-6๏ผŒwarmup 12 steps๏ผŒseed 20260901ใ€‚
  • SimPO๏ผš110 ็ป„ๅๅฅฝๅฏนใ€73 ไธชๅ”ฏไธ€ prompt๏ผŒ5 optimizer steps๏ผ›beta=1.0ใ€gamma=0.2ใ€peak LR 5e-7๏ผ›LoRA r=16ใ€alpha=32ใ€dropout=0๏ผŒseed 20260903๏ผŒworld size 2๏ผŒFSDP full shardingใ€‚
  • ่ฎญ็ปƒ็กฌไปถ๏ผšๅŒ NVIDIA RTX PRO 6000 Blackwell Server Editionใ€‚

ๆœฌไป“ๅ‘ๅธƒๆ–‡ๆœฌไธป GGUFใ€่ง†่ง‰ๆŠ•ๅฝฑไธŽ็ป่ฟ‡็‹ฌ็ซ‹้ชŒ่ฏ็š„ DFlash2 draftใ€‚Q8_0 ไธŽ Q6_K ไฝฟ็”จ Q8 ่ง†่ง‰ๅก”๏ผ›Q5-LynnStyleใ€Q4-LynnStyle ไธŽ Q3-LynnStyle ไฝฟ็”จ Q6 ่ง†่ง‰ๅก”ใ€‚ๅ…ญๆกฃ้ƒฝๆไพ› Q8_0 ไธŽๆ›ดๅฐ็š„ Q4_K_M DFlash2 draft๏ผŒๅนถๅฆๅซ Q4_0 ไธŽ Q8_0 MTP sidecarใ€‚ๆฏๆฌกๅฏๅŠจๅช้€‰ไธ€ๆกๆŠ•ๆœบ่งฃ็ ่ทฏ็บฟ๏ผ›่ฟ™ไบ› sidecar ้ƒฝไธๆ˜ฏไธปๆจกๅž‹ใ€‚

ๅทฒๅ‘ๅธƒ GGUF ๆ–‡ไปถ

Q8_0

ๆญฃๅผๅ…จ้‡่ฏ„ๆต‹๏ผšGPQA 164/198๏ผˆ82.83%๏ผ‰ยท MMLU 447/500๏ผˆ89.40%๏ผ‰ยท LCB 74/100๏ผˆ74.00%๏ผ‰ใ€‚

่ทฏๅพ„ ไฝœ็”จ ๅคงๅฐ
Q8_0/Qwen3.8-27B-EfficientThink-SimPO-Q8_0.gguf Q8_0 ๆ–‡ๆœฌไธปๆจกๅž‹ 28,595,763,552 bytes
Q8_0/mmproj-Qwen3.8-27B-Q8_0.gguf Q8 ่ง†่ง‰ๆŠ•ๅฝฑ 629,247,008 bytes
Q8_0/dflash2-qwen38-27b-Q8_0.gguf Q8_0 DFlash2 draft 2,056,414,720 bytes
Q8_0/dflash2-qwen38-27b-Q4_K_M.gguf ๅฏ้€‰ Q4_K_M DFlash2 draft๏ผ›ไธๆ˜ฏ MTP 1,143,006,720 bytes
Q8_0/mtp-Qwen3.8-27B-Q4_0.gguf ๅฏ้€‰ Q4_0 MTP draft๏ผ›ไธๆ˜ฏไธปๆจกๅž‹ๆˆ– DFlash2 1,680,271,648 bytes
Q8_0/mtp-Qwen3.8-27B-Q8_0.gguf ๅฏ้€‰ Q8_0 MTP draft๏ผ›ไธๆ˜ฏไธปๆจกๅž‹ๆˆ– DFlash2 3,164,006,688 bytes
Q8_0/SHA256SUMS ๆ–‡ไปถๆ ก้ชŒ 396 bytes
Q8_0/manifest.json ็ป“ๆž„ๅŒ–ๆธ…ๅ• 4,621 bytes

ๅนถๅ‘ๅปบ่ฎฎ๏ผšC4 ไธบๅฎž็”จๅนณ่กกๆกฃ๏ผˆ246 tok/s๏ผŒ60.98% ๆŽฅๅ—็އ๏ผ‰๏ผ›C8 ไธบๆœ€ๅคงๅฎžๆต‹ๅžๅ๏ผˆ256 tok/s๏ผŒ52.01%๏ผ‰ใ€‚

CUDA_VISIBLE_DEVICES=0 llama-server \
  --model Q8_0/Qwen3.8-27B-EfficientThink-SimPO-Q8_0.gguf \
  --mmproj Q8_0/mmproj-Qwen3.8-27B-Q8_0.gguf \
  --model-draft Q8_0/dflash2-qwen38-27b-Q8_0.gguf \
  --spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
  --gpu-layers all --gpu-layers-draft all \
  --parallel 8 --ctx-size 262144 --flash-attn on \
  --jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
  --host 127.0.0.1 --port 8080

Q5-LynnStyle

ๆญฃๅผๅ…จ้‡่ฏ„ๆต‹๏ผšGPQA 164/198๏ผˆ82.83%๏ผ‰ยท MMLU 438/500๏ผˆ87.60%๏ผ‰ยท LCB 75/100๏ผˆ75.00%๏ผ‰ใ€‚LCB๏ผšEasy 23/23ใ€Medium 29/31ใ€Hard 23/46ใ€‚

่ทฏๅพ„ ไฝœ็”จ ๅคงๅฐ
Q5-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q5-LynnStyle.gguf Q5-LynnStyle ๆ–‡ๆœฌไธปๆจกๅž‹ 22,412,452,608 bytes
Q5-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf Q6 ่ง†่ง‰ๆŠ•ๅฝฑ 618,390,560 bytes
Q5-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf Q8_0 DFlash2 draft 2,056,414,720 bytes
Q5-LynnStyle/dflash2-qwen38-27b-Q4_K_M.gguf ๅฏ้€‰ Q4_K_M DFlash2 draft๏ผ›ไธๆ˜ฏ MTP 1,143,006,720 bytes
Q5-LynnStyle/mtp-Qwen3.8-27B-Q4_0.gguf ๅฏ้€‰ Q4_0 MTP draft๏ผ›ไธๆ˜ฏไธปๆจกๅž‹ๆˆ– DFlash2 1,680,271,648 bytes
Q5-LynnStyle/mtp-Qwen3.8-27B-Q8_0.gguf ๅฏ้€‰ Q8_0 MTP draft๏ผ›ไธๆ˜ฏไธปๆจกๅž‹ๆˆ– DFlash2 3,164,006,688 bytes
Q5-LynnStyle/SHA256SUMS ๆ–‡ไปถๆ ก้ชŒ 501 bytes
Q5-LynnStyle/manifest.json ็ป“ๆž„ๅŒ–ๆธ…ๅ• 5,287 bytes

ๅนถๅ‘ๅปบ่ฎฎ๏ผšC4 ไธบๅฎž็”จๅนณ่กกๆกฃ๏ผˆ238 tok/s๏ผŒ48.22% ๆŽฅๅ—็އ๏ผ‰๏ผ›C8 ไธบๆœ€ๅคงๅฎžๆต‹ๅžๅ๏ผˆ276 tok/s๏ผŒ48.38%๏ผ‰ใ€‚

CUDA_VISIBLE_DEVICES=0 llama-server \
  --model Q5-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q5-LynnStyle.gguf \
  --mmproj Q5-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf \
  --model-draft Q5-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf \
  --spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
  --gpu-layers all --gpu-layers-draft all \
  --parallel 8 --ctx-size 262144 --flash-attn on \
  --jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
  --host 127.0.0.1 --port 8080

Q4-LynnStyle

ๆญฃๅผๅ…จ้‡่ฏ„ๆต‹๏ผšGPQA 166/198๏ผˆ83.84%๏ผ‰ยท MMLU 443/500๏ผˆ88.60%๏ผ‰ยท LCB 74/100๏ผˆ74.00%๏ผ‰ใ€‚LCB๏ผšEasy 23/23ใ€Medium 30/31ใ€Hard 21/46๏ผ›100 ไธชๅ”ฏไธ€้ข˜็›ฎใ€errors=0ใ€clean gate PASSใ€‚

่ทฏๅพ„ ไฝœ็”จ ๅคงๅฐ
Q4-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q4-LynnStyle.gguf Q4-LynnStyle ๆ–‡ๆœฌไธปๆจกๅž‹ 19,351,954,688 bytes
Q4-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf Q6 ่ง†่ง‰ๆŠ•ๅฝฑ 618,390,560 bytes
Q4-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf ๅ…ฑไบซ Q8_0 DFlash2 draft 2,056,414,720 bytes
Q4-LynnStyle/dflash2-qwen38-27b-Q4_K_M.gguf ๅฏ้€‰ Q4_K_M DFlash2 draft๏ผ›ไธๆ˜ฏ MTP 1,143,006,720 bytes
Q4-LynnStyle/mtp-Qwen3.8-27B-Q4_0.gguf ๅฏ้€‰ Q4_0 MTP draft๏ผ›ไธๆ˜ฏไธปๆจกๅž‹ๆˆ– DFlash2 1,680,271,648 bytes
Q4-LynnStyle/mtp-Qwen3.8-27B-Q8_0.gguf ๅฏ้€‰ Q8_0 MTP draft๏ผ›ไธๆ˜ฏไธปๆจกๅž‹ๆˆ– DFlash2 3,164,006,688 bytes
Q4-LynnStyle/SHA256SUMS ๆ–‡ไปถๆ ก้ชŒ 501 bytes
Q4-LynnStyle/manifest.json ็ป“ๆž„ๅŒ–ๆธ…ๅ• 5,439 bytes
ๅนถๅ‘ ่šๅˆๅžๅ DFlash2 ๆŽฅๅ—็އ ๅนณๅ‡ๆŽฅๅ—้•ฟๅบฆ / 8
C1 97.96 tok/s 38.83% 3.70
C2 149.27 tok/s 36.15% 3.51
C4 256.52 tok/s 47.69% 4.30
C8 263.50 tok/s 42.88% 3.95

ๅนถๅ‘ๅปบ่ฎฎ๏ผšC4 ไธบๅฎž็”จๅนณ่กกๆกฃ๏ผˆ257 tok/s๏ผŒ47.69% ๆŽฅๅ—็އ๏ผ‰๏ผ›C8 ไธบๆœ€ๅคงๅฎžๆต‹ๅžๅ๏ผˆ264 tok/s๏ผŒ42.88%๏ผ‰ใ€‚ๆ•ด่ฝฎๆ‰ซๅ‚ไธญ๏ผŒ่ถ…ๆ—ถใ€่ฏทๆฑ‚้”™่ฏฏใ€็ฉบ่พ“ๅ‡บใ€ๆ ‡็‚นๅๅกŒๅ‡ไธบ 0ใ€‚

CUDA_VISIBLE_DEVICES=0 llama-server \
  --model Q4-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q4-LynnStyle.gguf \
  --mmproj Q4-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf \
  --model-draft Q4-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf \
  --spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
  --gpu-layers all --gpu-layers-draft all \
  --parallel 8 --ctx-size 262144 --flash-attn on \
  --jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
  --host 127.0.0.1 --port 8080

Q3-LynnStyle

ๆญฃๅผๅ…จ้‡่ฏ„ๆต‹๏ผšGPQA 172/198๏ผˆ86.87%๏ผ‰ยท MMLU 435/500๏ผˆ87.00%๏ผ‰ยท LCB 78/100๏ผˆ78.00%๏ผ‰ใ€‚GPQA ไธบ 198 ไธชๅ”ฏไธ€้ข˜็›ฎใ€errors=0ใ€strict loop=0๏ผ›MMLU errors=0ใ€parse fail=0๏ผ›LCB๏ผšEasy 23/23ใ€Medium 28/31ใ€Hard 27/46๏ผŒ100 ไธชๅ”ฏไธ€้ข˜็›ฎใ€errors=0ใ€‚ๆ‰€ๆœ‰ๆœช้€š่ฟ‡ๆ ทๆœฌๅ‡ไฟ็•™ๅœจๅˆ†ๆฏใ€‚

่ทฏๅพ„ ไฝœ็”จ ๅคงๅฐ
Q3-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q3-LynnStyle.gguf Q3-LynnStyle ๆ–‡ๆœฌไธปๆจกๅž‹ 17,032,184,960 bytes
Q3-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf Q6 ่ง†่ง‰ๆŠ•ๅฝฑ 618,390,560 bytes
Q3-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf ๅ…ฑไบซ Q8_0 DFlash2 draft 2,056,414,720 bytes
Q3-LynnStyle/dflash2-qwen38-27b-Q4_K_M.gguf ๅฏ้€‰ Q4_K_M DFlash2 draft๏ผ›ไธๆ˜ฏ MTP 1,143,006,720 bytes
Q3-LynnStyle/mtp-Qwen3.8-27B-Q4_0.gguf ๅฏ้€‰ Q4_0 MTP draft๏ผ›ไธๆ˜ฏไธปๆจกๅž‹ๆˆ– DFlash2 1,680,271,648 bytes
Q3-LynnStyle/mtp-Qwen3.8-27B-Q8_0.gguf ๅฏ้€‰ Q8_0 MTP draft๏ผ›ไธๆ˜ฏไธปๆจกๅž‹ๆˆ– DFlash2 3,164,006,688 bytes
Q3-LynnStyle/SHA256SUMS ๆ–‡ไปถๆ ก้ชŒ 501 bytes
Q3-LynnStyle/manifest.json ็ป“ๆž„ๅŒ–ๆธ…ๅ• 5,526 bytes

ๅนถๅ‘ๅปบ่ฎฎ๏ผšC4 ไธบๅฎž็”จๅนณ่กกๆกฃ๏ผˆ269 tok/s๏ผŒ58.00% ๆŽฅๅ—็އ๏ผ‰๏ผ›C8 ไธบๆœ€ๅคงๅฎžๆต‹ๅžๅ๏ผˆ286 tok/s๏ผŒ52.41%๏ผ‰ใ€‚ๆ•ด่ฝฎๆ‰ซๅ‚ไธญ๏ผŒ่ถ…ๆ—ถใ€่ฏทๆฑ‚้”™่ฏฏใ€็ฉบ่พ“ๅ‡บใ€ๆ ‡็‚นๅๅกŒๅ‡ไธบ 0ใ€‚

CUDA_VISIBLE_DEVICES=0 llama-server \
  --model Q3-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q3-LynnStyle.gguf \
  --mmproj Q3-LynnStyle/mmproj-Qwen3.8-27B-Q6_K.gguf \
  --model-draft Q3-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf \
  --spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
  --gpu-layers all --gpu-layers-draft all \
  --parallel 8 --ctx-size 262144 --flash-attn on \
  --jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
  --host 127.0.0.1 --port 8080

Q2-LynnStyle

ๆญฃๅผๅ…จ้‡ๆˆ็ปฉ๏ผšGPQA 167/198๏ผˆ84.34%๏ผ‰ยท MMLU 416/500๏ผˆ83.20%๏ผ‰ยท LCB 75/100๏ผˆ75.00%๏ผ‰ใ€‚LCB๏ผšEasy 23/23ใ€Medium 29/31ใ€Hard 23/46๏ผ›23 ไธช 32K ้•ฟๅบฆ็ป“ๆŸๅ‡ไปฅๅคฑ่ดฅ่ฎกๅ…ฅใ€‚

LCB ๆœ€็ปˆ่ง†ๅ›พ็”ฑๅŽŸๅง‹่ฟ่กŒไธญไฟๆŒ้€ๅญ—่Š‚ไธๅ˜็š„ 99 ้ข˜๏ผŒๅŠ ไธŠ Lynn ๆŽˆๆƒๅฏน atcoder:arc181_b:81 ็š„ไธ€ๆฌก็ฒพ็กฎ่กฅๆต‹็ป„ๆˆใ€‚่ฏฅ้ข˜ๅŽŸ่ฟ่กŒๅ‘็”Ÿๆตๅผ JSON ้”™่ฏฏ๏ผ›่กฅๆต‹ไฝฟ็”จๅŒไธ€ๆจกๅž‹ใ€runnerใ€่ฏทๆฑ‚ไธŽๅ‚ๆ•ฐ๏ผŒๅฎŒๆ•ดๆ”ถๅˆฐ DONE๏ผŒไฝ†ไปๅœจ 32K ็ป“ๆŸไธ”ๆฒกๆœ‰ๆไบคไปฃ็ ๏ผŒๅ› ๆญคๆ€ปๅˆ†ไปไธบ 75/100ใ€‚่ฏฅ็ป“ๆžœไธๆ˜ฏไธ€ๆฌก่ฟž็ปญๆ— ๆ•…้šœ่ฟ่กŒใ€‚

่ทฏๅพ„ ไฝœ็”จ ๅคงๅฐ
Q2-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q2-LynnStyle.gguf Q2-LynnStyle ๆ–‡ๆœฌไธปๆจกๅž‹ 12,999,977,600 bytes
Q2-LynnStyle/mmproj-Qwen3.8-27B-Q4_K_M.gguf Q4_K_M ่ง†่ง‰ๆŠ•ๅฝฑ 522,293,024 bytes
Q2-LynnStyle/dflash2-qwen38-27b-Q8_0.gguf ๅฏ้€‰ Q8_0 DFlash2 draft๏ผ›ไธๆ˜ฏ MTP 2,056,414,720 bytes
Q2-LynnStyle/dflash2-qwen38-27b-Q4_K_M.gguf ๅฏ้€‰ Q4_K_M DFlash2 draft๏ผ›ไธๆ˜ฏ MTP 1,143,006,720 bytes
Q2-LynnStyle/mtp-Qwen3.8-27B-Q4_0.gguf ๅฏ้€‰ Q4_0 MTP draft๏ผ›ไธๆ˜ฏไธปๆจกๅž‹ๆˆ– DFlash2 1,680,271,648 bytes
Q2-LynnStyle/mtp-Qwen3.8-27B-Q8_0.gguf ๅฏ้€‰ Q8_0 MTP draft๏ผ›ไธๆ˜ฏไธปๆจกๅž‹ๆˆ– DFlash2 3,164,006,688 bytes
Q2-LynnStyle/SHA256SUMS ๆ–‡ไปถๆ ก้ชŒ 406 bytes
Q2-LynnStyle/manifest.json ็ป“ๆž„ๅŒ–ๆธ…ๅ• 5,504 bytes

Q2-LynnStyle ไฝฟ็”จ GSQ-RCO ๆททๅˆ็ฒพๅบฆ้‡ๅŒ–๏ผŒๅนถไปฅ IQ ๆ•ฐๅ€ผ็ป†ๅŒ–ไผ˜ๅŒ– 70 ไธช้€‰ๅฎš MLP ็Ÿฉ้˜ต๏ผ›attentionใ€head ไธŽ SSM tensor ๆœชๅšๆ•ฐๅ€ผ็ป†ๅŒ–ใ€‚่‹ฅๅ†…ๅญ˜ๅ…่ฎธ๏ผŒQ3-LynnStyle ๅŠๆ›ด้ซ˜ๆกฃไฝไปๆ˜ฏๆ›ด็จณๅฆฅ็š„่ดจ้‡้€‰ๆ‹ฉใ€‚ๆญค็ฒพ็กฎ 12,999,977,600-byte ๆž„ๅปบๆฒกๆœ‰ๅ†ป็ป“ TPS๏ผŒไธ่ƒฝๅฅ—็”จๅทฒๆท˜ๆฑฐๅทฅ็จ‹ Q2 ็š„ๅžๅใ€‚

CUDA_VISIBLE_DEVICES=0 llama-server \
  --model Q2-LynnStyle/Qwen3.8-27B-EfficientThink-SimPO-Q2-LynnStyle.gguf \
  --mmproj Q2-LynnStyle/mmproj-Qwen3.8-27B-Q4_K_M.gguf \
  --model-draft Q2-LynnStyle/dflash2-qwen38-27b-Q4_K_M.gguf \
  --spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
  --gpu-layers all --gpu-layers-draft all \
  --parallel 4 --ctx-size 131072 --flash-attn on \
  --jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
  --host 127.0.0.1 --port 8080

Q8_0ใ€Q6_Kใ€Q5-LynnStyleใ€Q4-LynnStyleใ€Q3-LynnStyle ไธŽ Q2-LynnStyle ็š„ไธป GGUF ๅ‡ไธบ GGUF v3ใ€qwen35ใ€64 ๅฑ‚ใ€851 tensors๏ผŒไธปๆ–‡ไปถไธๅ†…ๅตŒ่ง†่ง‰ๆˆ– MTP tensor๏ผ›ๅ‡ๅทฒ้€š่ฟ‡่ฟœ็ซฏๅญ—่Š‚/SHA256 ๅ›ž่ฏปใ€‚ๅฝ“ๅ‰ๅ‘ๅธƒ็š„ Q8_0 DFlash2 ๅทฒ้€š่ฟ‡้™ๆ€่งฃๆžไธŽ Q2-LynnStyle ไปฃ่กจๆ€ง่ฟ่กŒ smoke๏ผ›ๅ…ถไฝ™ๅˆ†ๆกฃๅฎŒๆ•ดๅนถๅ‘้‡ๆต‹ๅพ…่กฅ๏ผŒๆ—ง DFlash2 ๆ•ฐๆฎไธๅฝ’ๅ› ไบŽ่ฏฅ Q8_0 ๆ–‡ไปถใ€‚

LynnStyle ๆททๅˆ็ฒพๅบฆ้‡ๅŒ–

Q5/Q4/Q3 LynnStyle ๅ‡ไปŽๆœ€็ปˆ SimPO F16 ๆ–‡ๆœฌไธปๅนฒ็›ดๆŽฅ้‡ๅŒ–๏ผŒไฝฟ็”จๅ†ป็ป“็š„ 512-chunk imatrix๏ผˆ496 entries๏ผ‰๏ผŒไธๅšไบŒๆฌก requantizeใ€‚ไฟๆŠค็ญ–็•ฅ็ปผๅˆ T10 REAPใ€้ฆ–ๅฐพๅฑ‚/SSM ไฟๆŠคไธŽ full-attentionใ€SSMใ€FFN ็š„้€ tensor ๅฎก่ฎก๏ผ›่ง„ๅˆ™ๅ†ฒ็ชๆ—ถ้‡‡็”จๆ›ด้ซ˜็ฒพๅบฆใ€‚

  • Q5-LynnStyle๏ผˆๅทฒๅ‘ๅธƒ๏ผ‰๏ผšๅŸบ็ก€ Q5๏ผ›L0โ€“3ใ€L60โ€“63 ไธŽๅ…ณ้”ฎ REAP ๅฑ‚ 50/52/54/56/58 ไฝฟ็”จ Q8๏ผ›้ซ˜ๆ•ๆ„Ÿๅฑ‚ 21/42/44/46/48/49/53/57/59 ไฝฟ็”จ Q6๏ผ›full-attention q/k/v/o ไธŽ้€‰ๅฎšไธญๅฑ‚ ssm_out ไฝฟ็”จ Q6ใ€‚498 ไธช่ฝฌๆข tensor๏ผšQ8_0 104ใ€Q6_K 186ใ€Q5_K 208ใ€‚
  • Q4-LynnStyle๏ผˆๅทฒๅ‘ๅธƒ๏ผ‰๏ผšๅŸบ็ก€ Q4๏ผ›L0โ€“3ใ€L60โ€“63 ไฝฟ็”จ Q8๏ผ›ๅ…ณ้”ฎ่‚ฉ้ƒจๅฑ‚ 50/52/54/56โ€“59 ไฝฟ็”จ Q6๏ผ›้ซ˜ๆ•ๆ„Ÿๅฑ‚ 21/42/44/46/48/49/53 ไฝฟ็”จ Q5๏ผ›full-attention q/k/v/o ไธŽ้€‰ๅฎš ssm_out ไฝฟ็”จ Q6ใ€‚ๅˆ†ๅธƒ๏ผšQ8 62ใ€Q6 156ใ€Q5 172ใ€Q4 108ใ€‚
  • Q3-LynnStyle๏ผˆๅทฒๅ‘ๅธƒ๏ผ‰๏ผšๅŸบ็ก€ Q3๏ผ›L0โ€“3ใ€L60โ€“63 ไฝฟ็”จ Q6๏ผ›L4โ€“6ใ€50/52/54/56โ€“59 ไฝฟ็”จ Q5๏ผ›้ซ˜ๆ•ๆ„Ÿๅฑ‚ 21/42/44/46/48/49/53 ไฝฟ็”จ Q4๏ผ›full-attention ไธŽไธญๅฑ‚ ssm_out ไฝฟ็”จ Q5ใ€‚ๅˆ†ๅธƒ๏ผšQ8 2ใ€Q6 62ใ€Q5 157ใ€Q4 199ใ€Q3 78ใ€‚

Hub ไพง่พนๆ ๅฏ่ƒฝไป…ๆ˜พ็คบ GGUF ๅไน‰ general.file_type๏ผŒๆ— ๆณ•ๅฎŒๆ•ด่กจ่พพ่‡ชๅฎšไน‰้€ tensor ๆททๅˆ็ฒพๅบฆ๏ผ›ไปฅๅ„็›ฎๅฝ• manifest.jsonใ€SHA256SUMS ไธŽๅฎž้™… GGUF parse ไธบๅ‡†ใ€‚

Q6_K

ๆญฃๅผๅ…จ้‡่ฏ„ๆต‹๏ผšGPQA 171/198๏ผˆ86.36%๏ผ‰ยท MMLU 440/500๏ผˆ88.00%๏ผ‰ยท LCB 78/100๏ผˆ78.00%๏ผ‰ใ€‚

LCB ้šพๅบฆๅˆ†ๅธƒ๏ผšEasy 23/23ใ€Medium 30/31ใ€Hard 25/46ใ€‚

ๅนถๅ‘ ่šๅˆๅžๅ DFlash2 ๆŽฅๅ—็އ ๅนณๅ‡ๆŽฅๅ—้•ฟๅบฆ / 8 ๆœ€ๅคง่ฏทๆฑ‚ๆ—ถ้•ฟ
C1 91 tok/s 40.83% 3.79 2.80s
C2 165 tok/s 48.28% 4.36 3.10s
C4๏ผˆไฝŽๅปถ่ฟŸๅนณ่กกๆกฃ๏ผ‰ 242 tok/s 49.00% 4.37 4.23s
C8๏ผˆๆœ€ๅคงๅžๅ๏ผ‰ 280 tok/s 46.80% 4.24 7.31s

ๅ››ๆกฃๅ›บๅฎš่พ“ๅ‡บๆต‹่ฏ•็š„่ฏทๆฑ‚้”™่ฏฏใ€่ถ…ๆ—ถใ€็ฉบ่พ“ๅ‡บๅ’Œๆ ‡็‚นๅๅกŒๅ‡ไธบ 0ใ€‚ๅžๅๆต‹่ฏ•ไธไปฃๆ›ฟๆญฃๅผ่ดจ้‡่ฏ„ๆต‹ใ€‚

ๆ–‡ไปถ

่ทฏๅพ„ ไฝœ็”จ ๅคงๅฐ
Q6_K/Qwen3.8-27B-EfficientThink-SimPO-Q6_K.gguf Q6_K ๆ–‡ๆœฌไธปๆจกๅž‹ 22,082,530,048 bytes
Q6_K/mmproj-Qwen3.8-27B-Q8_0.gguf Q8 ่ง†่ง‰ๆŠ•ๅฝฑ 629,247,008 bytes
Q6_K/dflash2-qwen38-27b-Q8_0.gguf Q8_0 DFlash2 draft 2,056,414,720 bytes
Q6_K/dflash2-qwen38-27b-Q4_K_M.gguf ๅฏ้€‰ Q4_K_M DFlash2 draft๏ผ›ไธๆ˜ฏ MTP 1,143,006,720 bytes
Q6_K/mtp-Qwen3.8-27B-Q4_0.gguf ๅฏ้€‰ Q4_0 MTP draft๏ผ›ไธๆ˜ฏไธปๆจกๅž‹ๆˆ– DFlash2 1,680,271,648 bytes
Q6_K/mtp-Qwen3.8-27B-Q8_0.gguf ๅฏ้€‰ Q8_0 MTP draft๏ผ›ไธๆ˜ฏไธปๆจกๅž‹ๆˆ– DFlash2 3,164,006,688 bytes

ๆŽจ่ๅฏๅŠจ

CUDA_VISIBLE_DEVICES=0 llama-server \
  --model Q6_K/Qwen3.8-27B-EfficientThink-SimPO-Q6_K.gguf \
  --mmproj Q6_K/mmproj-Qwen3.8-27B-Q8_0.gguf \
  --model-draft Q6_K/dflash2-qwen38-27b-Q8_0.gguf \
  --spec-type draft-dflash --spec-draft-n-max 7 --spec-draft-n-min 0 \
  --gpu-layers all --gpu-layers-draft all \
  --parallel 8 --ctx-size 262144 --flash-attn on \
  --jinja --chat-template-kwargs '{"enable_thinking":true,"reasoning_effort":"xhigh"}' \
  --host 127.0.0.1 --port 8080
Downloads last month
87,129
GGUF
Model size
2B params
Architecture
dflash
Hardware compatibility
Log In to add your hardware

4-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF

Quantizations
1 model

Space using nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-GGUF 1