Qwen-Image-2.1-PE-T2I-Pocket-0.8B

Built with Qwen — non-commercial research use (Qwen Research License, see LICENSE in this repo)

A tiny text-only prompt rewriter for Qwen/Qwen-Image-2.1: it turns a short user request ("a photo of a red bicycle leaning against a bakery door") into the long English description and aspect ratio that the image model expects from its prompt-rewriting step — without the 9B thinking model Qwen/Qwen-Image-2.1-PE-T2I, its 1,700-word system prompt, or any thinking tokens.

It is a full fine-tune of Qwen/Qwen3.5-0.8B (Apache 2.0, no architectural change), distilled from teacher outputs: {user: raw request} → {assistant: one compact JSON object}. No system prompt and thinking are disabled in the chat template, so the rewrite behaviour — including the JSON format — is baked into the weights.

{"rewritten_prompt": "...", "wh_ratio": "3:2"}

Teacher, data and filter flags: ML-Intern-lab/Qwen-Image-2.1-rewriter-distill. Training data is the hard-filtered subset (1,776 pairs) of 8,797 teacher-labelled requests: parse_ok, allowed ratio, quoted text verbatim, user-stated ratio respected, English output, 80–400 words, no ratio/resolution/pixel text inside the prompt.

Usage

import json
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from diffusers import QwenImage21Pipeline  # diffusers from git main: 0.41.0.dev0+

tok = AutoTokenizer.from_pretrained("ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B")
model = AutoModelForCausalLM.from_pretrained(
    "ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B",
    torch_dtype=torch.bfloat16,
    device_map="cuda",
)

def rewrite(request: str) -> dict:
    """Raw request -> {"rewritten_prompt": str, "wh_ratio": str}.

    No system prompt, thinking disabled: the chat template emits an empty
    think block before the answer, exactly as at training time, so the model
    continues directly with the JSON object.
    """
    prompt = tok.apply_chat_template(
        [{"role": "user", "content": request}],
        add_generation_prompt=True, tokenize=False,
    )
    ids = tok(prompt, return_tensors="pt").to(model.device)
    gen = model.generate(
        **ids,
        max_new_tokens=1024,
        do_sample=True, temperature=1.0, top_p=0.95, top_k=20,
        pad_token_id=tok.pad_token_id or tok.eos_token_id,
    )
    text = tok.decode(gen[0][ids["input_ids"].shape[1]:], skip_special_tokens=True)
    return json.loads(text)  # rare ~0.3% malformed: resample or fall back to json_repair

# Render with Qwen-Image-2.1 (diffusers from git main; torchvision must be importable)
pipe = QwenImage21Pipeline.from_pretrained(
    "Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16, device_map="cuda"
)

# ~1 megapixel (height, width), multiples of 32 — same table as the image eval
RATIO_SIZES = {
    "1:1": (1024, 1024), "3:2": (1248, 832), "2:3": (832, 1248),
    "16:9": (1376, 768), "9:16": (768, 1376), "4:3": (1184, 896),
    "3:4": (896, 1184), "2:1": (1472, 736), "1:2": (736, 1472),
    "21:9": (1568, 672), "9:21": (672, 1568), "4:5": (928, 1152),
    "5:4": (1152, 928), "3:1": (1728, 576), "1:3": (576, 1760),
}

r = rewrite("a photo of a red bicycle leaning against a bakery door")
w, h = RATIO_SIZES.get(r["wh_ratio"], (1248, 832))
image = pipe(
    r["rewritten_prompt"], height=h, width=w,
    num_inference_steps=40,          # default 40 steps, true_cfg_scale 1.0 = no CFG
).images[0]
image.save("bicycle.png")

Evaluation

Measured on the 300 held-out eval requests of the distill dataset, generating every arm with the teacher's own sampling protocol (temperature 1.0, top_p 0.95, top_k 20, seed 0). Baseline = the untuned Qwen/Qwen3.5-2B given the teacher's full system prompt with thinking enabled (its run was cut at 80 rows by the eval job's time budget, so its numbers are a partial floor, not a ceiling).

metric teacher (9B) this model student-2B baseline-2B
rows 300 300 300 80
valid JSON rate 100.0% 99.7% 100.0% 77.5%
allowed-ratio rate 100.0% 99.3% 99.7% 20.0%
text fidelity (quoted strings verbatim) 53.1% 53.1% 60.2% 3.7%
fidelity by script Arabic 66.7%, Devanagari 0.0%, Han 60.0%, Japanese 66.7%, Latin 52.0% Arabic 33.3%, Devanagari 0.0%, Han 20.0%, Japanese 33.3%, Latin 57.1% Arabic 33.3%, Devanagari 0.0%, Han 20.0%, Japanese 50.0%, Latin 64.3% Japanese 0.0%, Latin 3.9%
ratio agreement w/ teacher 100.0% 57.7% 67.3% 8.8%
gen tokens mean / median 1630.8 / 1536.5 453.1 / 456.5 482.8 / 462.0 6106.4 / 6144.0
latency s mean / median n/a 2.81 / 2.61 3.22 / 3.12 29.90 / 29.81

Reads:

  • The students match the teacher on format compliance (JSON + allowed ratio ≈ 99–100%) at ~28% of the teacher's token cost and a fraction of the latency.
  • The teacher itself only preserves ~53% of quoted in-image text; the students are in the same band, and all three lose more on Arabic/Devanagari than on Latin. This is a property of the whole stack, not of the students alone.
  • Ratio: students nearly always choose an allowed ratio but match the teacher's specific choice only 58–67% of the time.

GGUF Q8 for CPU

Qwen-Image-2.1-PE-T2I-Pocket-0.8B-Q8_0.gguf (812 MB) — llama.cpp Q8_0 quantization of this model, converted with convert_hf_to_gguf.py --no-nextn (the config's MTP block is architecture-only; it has no weights) and quantized with llama-quantize. Needs a recent llama.cpp build with Qwen3.5 linear-attention support.

llama-cli -m Qwen-Image-2.1-PE-T2I-Pocket-0.8B-Q8_0.gguf \
  -p '<|im_start|>user
a photo of a red bicycle leaning against a bakery door<|im_end|>
<|im_start|>assistant
' -n 1024 --temp 1.0 --top-p 0.95 --top-k 20 -no-cnv

Measured on a Xeon 8375C (12 threads): ~6.85 tokens/s → ~66 s per rewrite at the student's 453-token mean; sampled completions parse as valid JSON.

Limitations

  • Non-commercial. Trained on Qwen-Image-2.1-PE-T2I outputs under the Qwen Research License: research/evaluation use only; commercial use requires a separate licence from Qwen (see NOTICE).
  • English output for any input. Like the teacher, it rewrites non-English requests into English descriptions; quoted in-image text is copied verbatim in the original script at roughly the teacher's rate (see fidelity by script — weaker for Arabic/Devanagari than Latin).
  • Shorter than the teacher. Trained on rewrites of 80–400 words (fits max_length 1024); the teacher writes up to ~4× longer. The student compresses — fine for generation, but not a drop-in for the teacher when extreme detail is wanted.
  • Format is learned, not guaranteed. ~0.3% of generations are not valid JSON; ~0.7% of ratios fall outside the allowed set. Resample or repair rather than trusting any single draw.
  • Ratio choice. A user-stated ratio was enforced in training, but the model does not always follow it, and it often picks a different reasonable ratio than the teacher would.
  • Not a general chat model. Single task; baked-in instructions, no system prompt. Give it an image request, not a conversation.
  • Image-level evaluation pending. The text metrics above are necessary but not sufficient; the OCR + VLM-judge render eval (40 requests × 4 arms) lives in the dataset repo under eval/ and renders/sheets/.

Training

TRL 1.13.0 SFTTrainer, full fine-tune, bf16, 2 epochs, lr 1e-5 cosine (warmup 3%), effective batch 32 (8×4), max_length 1024, assistant-only loss, gradient checkpointing, seed 42. Final train loss 1.693 (plateaued — see the run log for the shape). Trackio: https://huggingface.co/spaces/ML-Intern-lab/Qwen-Image-2.1-rewriter-distill-trackio

License

This model is a derivative of Qwen-Image-2.1-PE-T2I outputs and is distributed under the Qwen Research License Agreement — see LICENSE (full text) and NOTICE (attribution per §3(c), modification statement per §3(b), and the Built with Qwen mark per §4(b)). Non-commercial use only.

Downloads last month
4,350
Safetensors
Model size
0.8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B

Quantized
(320)
this model
Quantizations
1 model

Spaces using ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B 4

Collection including ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B