clef-flash-oQ4e

Updated 2026-10-11: Re-quantized with Improved oQe Quantization. Please re-download if you have an earlier copy.

This model was quantized using oQ (oMLX v0.7.1) mixed-precision quantization with Improved oQe Quantization (jundot/omlx#4385).

A mixed-precision MLX checkpoint of Cloudflare/clef-flash, a 9B decision model post-trained from Qwen/Qwen3.5-9B that turns a state and a schema of typed questions into decisions, returning a probability for every allowed option of every question in a single forward pass. It includes the Qwen3.5-9B language backbone, the vision encoder and the joint schema head.

The joint schema head (joint_head.safetensors, joint_head_config.json) is copied unchanged from the source. oMLX serves this checkpoint through POST /v1/systemone.

The effective average is 5.268 bits/weight for the complete checkpoint.

Quantization and Bit Distribution

oQ4e allocates bits by measured layer sensitivity, then rounds each group with an importance matrix (128 samples x 512 tokens) and a weighted least-squares refit of its scale and bias.

Storage format Logical weights Share of total Tensor storage Effective bits/weight
Affine 4-bit, group 64 8.288B 86.948% 4.342 GiB 4.50
Affine 5-bit, group 64 0.665B 6.979% 0.426 GiB 5.50
BF16 0.579B 6.073% 1.078 GiB 16.00
Total 9.532B 100% 5.846 GiB 5.268

Storage Breakdown

Component Logical weights Storage
Language backbone 8.954B 5.121 GB / 4.769 GiB
Vision encoder 0.456B 0.912 GB / 0.849 GiB
Joint schema head 0.122B 0.244 GB / 0.227 GiB
Total 9.532B 6.277 GB / 5.846 GiB

License

Apache-2.0, following the source model. See LICENSE.

Downloads last month
415
Safetensors
Model size
9B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Jundot/clef-flash-oQ4e

Finetuned
Qwen/Qwen3.5-9B
Quantized
(53)
this model