ThinkingCap Qwen 3.8 Quants
Quantized ThinkingCap-Qwen3.8-27B: NVFP4 and NVFP4A4-AWQ (vLLM), GGUF (llama.cpp), MLX (Apple silicon).
Image-Text-to-Text âĒ 17B âĒ Updated âĒ 2.3k âĒ 27Note ð NVFP4 weight-only, tied GDN scales, vLLM, 20.6 GB. Hopper (Marlin) and Blackwell; needs a 32 GB GPU (RTX 5090). Evals vs bf16: accuracy â3.0 to +0.9 pp per benchmark. Mean stronger shortening, median kept.
bottlecapai/ThinkingCap-Qwen3.8-27B-NVFP4A4-AWQ
Image-Text-to-Text âĒ 20B âĒ Updated âĒ 1.49k âĒ 12Note ðĪ NVFP4 weights and activations (AWQ), vLLM, 23.4 GB. BLACKWELL ONLY (CUTLASS FP4); 32 GB GPU. Meant for batched serving over NVFP4. Evals vs bf16: accuracy â1.3 to +1.0 pp, shortening kept.
bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF
Image-Text-to-Text âĒ 27B âĒ Updated âĒ 455k âĒ 79Note ðĶ GGUF, llama.cpp/LM Studio/Ollama; CUDA, Metal, Vulkan or CPU. Weights plus ~2 GB per 32k context (+0.9 GB mmproj): ðĶ IQ4_XS 15.5 GB, ð Q4_K_M 17.4 GB (24 GB GPU, 32 GB Mac); ð Q6_K 23.9 (32 GB GPU), ð Q8_0 29.0 (48 GB GPU); ðŧ f16 54.7 (80 GB). Q4_K_M evals vs bf16: accuracy â3.0 to +0.8 pp per benchmark, none significant. Shortening kept.
bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF-RCO-4.2bpw
Image-Text-to-Text âĒ 27B âĒ Updated âĒ 2.59k âĒ 4Note ðŠķ GGUF with per-tensor mixed precision (RCO search), 4.19 bits per weight, 14.3 GB (+0.9 GB mmproj); llama.cpp/LM Studio/Ollama. Fits an 18 GB GPU with vision, MTP and 32k context (q8_0 KV; 40k still fits). Evals vs bf16: accuracy â0.7 to +1.0 pp per benchmark, none significant. Shortening kept.
bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ
Image-Text-to-Text âĒ 27B âĒ Updated âĒ 2.73k âĒ 16Note ð MLX for Apple silicon, mixed 4/8-bit DWQ, 22.5 GB; runs on a 32 GB Mac. Serve with mlx-vlm or oMLX: vision input and MTP self-speculative decoding work there (mlx-lm drops both).