Instructions to use ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B") model = AutoModelForCausalLM.from_pretrained("ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B:Q8_0 # Run inference directly in the terminal: llama cli -hf ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B:Q8_0 # Run inference directly in the terminal: llama cli -hf ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B:Q8_0
Use Docker
docker model run hf.co/ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B:Q8_0
- LM Studio
- Jan
- vLLM
How to use ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B:Q8_0
- SGLang
How to use ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B with Ollama:
ollama run hf.co/ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B:Q8_0
- Unsloth Desktop
- Pi
How to use ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B with Docker Model Runner:
docker model run hf.co/ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B:Q8_0
- Lemonade
How to use ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B:Q8_0
Run and chat with the model
lemonade run user.Qwen-Image-2.1-PE-T2I-Pocket-0.8B-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen-Image-2.1-PE-T2I-Pocket-0.8B
Built with Qwen — non-commercial research use (Qwen Research License, see LICENSE in this repo)
A tiny text-only prompt rewriter for Qwen/Qwen-Image-2.1: it turns a short user request ("a photo of a red bicycle leaning against a bakery door") into the long English description and aspect ratio that the image model expects from its prompt-rewriting step — without the 9B thinking model Qwen/Qwen-Image-2.1-PE-T2I, its 1,700-word system prompt, or any thinking tokens.
It is a full fine-tune of Qwen/Qwen3.5-0.8B (Apache 2.0, no architectural change), distilled from teacher outputs: {user: raw request} → {assistant: one compact JSON object}. No system prompt and thinking are disabled in the chat template, so the rewrite behaviour — including the JSON format — is baked into the weights.
{"rewritten_prompt": "...", "wh_ratio": "3:2"}
Teacher, data and filter flags: ML-Intern-lab/Qwen-Image-2.1-rewriter-distill. Training data is the hard-filtered subset (1,776 pairs) of 8,797 teacher-labelled requests: parse_ok, allowed ratio, quoted text verbatim, user-stated ratio respected, English output, 80–400 words, no ratio/resolution/pixel text inside the prompt.
Usage
import json
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from diffusers import QwenImage21Pipeline # diffusers from git main: 0.41.0.dev0+
tok = AutoTokenizer.from_pretrained("ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B")
model = AutoModelForCausalLM.from_pretrained(
"ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B",
torch_dtype=torch.bfloat16,
device_map="cuda",
)
def rewrite(request: str) -> dict:
"""Raw request -> {"rewritten_prompt": str, "wh_ratio": str}.
No system prompt, thinking disabled: the chat template emits an empty
think block before the answer, exactly as at training time, so the model
continues directly with the JSON object.
"""
prompt = tok.apply_chat_template(
[{"role": "user", "content": request}],
add_generation_prompt=True, tokenize=False,
)
ids = tok(prompt, return_tensors="pt").to(model.device)
gen = model.generate(
**ids,
max_new_tokens=1024,
do_sample=True, temperature=1.0, top_p=0.95, top_k=20,
pad_token_id=tok.pad_token_id or tok.eos_token_id,
)
text = tok.decode(gen[0][ids["input_ids"].shape[1]:], skip_special_tokens=True)
return json.loads(text) # rare ~0.3% malformed: resample or fall back to json_repair
# Render with Qwen-Image-2.1 (diffusers from git main; torchvision must be importable)
pipe = QwenImage21Pipeline.from_pretrained(
"Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16, device_map="cuda"
)
# ~1 megapixel (height, width), multiples of 32 — same table as the image eval
RATIO_SIZES = {
"1:1": (1024, 1024), "3:2": (1248, 832), "2:3": (832, 1248),
"16:9": (1376, 768), "9:16": (768, 1376), "4:3": (1184, 896),
"3:4": (896, 1184), "2:1": (1472, 736), "1:2": (736, 1472),
"21:9": (1568, 672), "9:21": (672, 1568), "4:5": (928, 1152),
"5:4": (1152, 928), "3:1": (1728, 576), "1:3": (576, 1760),
}
r = rewrite("a photo of a red bicycle leaning against a bakery door")
w, h = RATIO_SIZES.get(r["wh_ratio"], (1248, 832))
image = pipe(
r["rewritten_prompt"], height=h, width=w,
num_inference_steps=40, # default 40 steps, true_cfg_scale 1.0 = no CFG
).images[0]
image.save("bicycle.png")
Evaluation
Measured on the 300 held-out eval requests of the distill dataset, generating every arm with the teacher's own sampling protocol (temperature 1.0, top_p 0.95, top_k 20, seed 0). Baseline = the untuned Qwen/Qwen3.5-2B given the teacher's full system prompt with thinking enabled (its run was cut at 80 rows by the eval job's time budget, so its numbers are a partial floor, not a ceiling).
| metric | teacher (9B) | this model | student-2B | baseline-2B |
|---|---|---|---|---|
| rows | 300 | 300 | 300 | 80 |
| valid JSON rate | 100.0% | 99.7% | 100.0% | 77.5% |
| allowed-ratio rate | 100.0% | 99.3% | 99.7% | 20.0% |
| text fidelity (quoted strings verbatim) | 53.1% | 53.1% | 60.2% | 3.7% |
| fidelity by script | Arabic 66.7%, Devanagari 0.0%, Han 60.0%, Japanese 66.7%, Latin 52.0% | Arabic 33.3%, Devanagari 0.0%, Han 20.0%, Japanese 33.3%, Latin 57.1% | Arabic 33.3%, Devanagari 0.0%, Han 20.0%, Japanese 50.0%, Latin 64.3% | Japanese 0.0%, Latin 3.9% |
| ratio agreement w/ teacher | 100.0% | 57.7% | 67.3% | 8.8% |
| gen tokens mean / median | 1630.8 / 1536.5 | 453.1 / 456.5 | 482.8 / 462.0 | 6106.4 / 6144.0 |
| latency s mean / median | n/a | 2.81 / 2.61 | 3.22 / 3.12 | 29.90 / 29.81 |
Reads:
- The students match the teacher on format compliance (JSON + allowed ratio ≈ 99–100%) at ~28% of the teacher's token cost and a fraction of the latency.
- The teacher itself only preserves ~53% of quoted in-image text; the students are in the same band, and all three lose more on Arabic/Devanagari than on Latin. This is a property of the whole stack, not of the students alone.
- Ratio: students nearly always choose an allowed ratio but match the teacher's specific choice only 58–67% of the time.
GGUF Q8 for CPU
Qwen-Image-2.1-PE-T2I-Pocket-0.8B-Q8_0.gguf (812 MB) — llama.cpp Q8_0 quantization of this model, converted with convert_hf_to_gguf.py --no-nextn (the config's MTP block is architecture-only; it has no weights) and quantized with llama-quantize. Needs a recent llama.cpp build with Qwen3.5 linear-attention support.
llama-cli -m Qwen-Image-2.1-PE-T2I-Pocket-0.8B-Q8_0.gguf \
-p '<|im_start|>user
a photo of a red bicycle leaning against a bakery door<|im_end|>
<|im_start|>assistant
' -n 1024 --temp 1.0 --top-p 0.95 --top-k 20 -no-cnv
Measured on a Xeon 8375C (12 threads): ~6.85 tokens/s → ~66 s per rewrite at the student's 453-token mean; sampled completions parse as valid JSON.
Limitations
- Non-commercial. Trained on Qwen-Image-2.1-PE-T2I outputs under the Qwen Research License: research/evaluation use only; commercial use requires a separate licence from Qwen (see NOTICE).
- English output for any input. Like the teacher, it rewrites non-English requests into English descriptions; quoted in-image text is copied verbatim in the original script at roughly the teacher's rate (see fidelity by script — weaker for Arabic/Devanagari than Latin).
- Shorter than the teacher. Trained on rewrites of 80–400 words (fits
max_length 1024); the teacher writes up to ~4× longer. The student compresses — fine for generation, but not a drop-in for the teacher when extreme detail is wanted. - Format is learned, not guaranteed. ~0.3% of generations are not valid JSON; ~0.7% of ratios fall outside the allowed set. Resample or repair rather than trusting any single draw.
- Ratio choice. A user-stated ratio was enforced in training, but the model does not always follow it, and it often picks a different reasonable ratio than the teacher would.
- Not a general chat model. Single task; baked-in instructions, no system prompt. Give it an image request, not a conversation.
- Image-level evaluation pending. The text metrics above are necessary but not sufficient; the OCR + VLM-judge render eval (40 requests × 4 arms) lives in the dataset repo under
eval/andrenders/sheets/.
Training
TRL 1.13.0 SFTTrainer, full fine-tune, bf16, 2 epochs, lr 1e-5 cosine (warmup 3%), effective batch 32 (8×4), max_length 1024, assistant-only loss, gradient checkpointing, seed 42. Final train loss 1.693 (plateaued — see the run log for the shape). Trackio: https://huggingface.co/spaces/ML-Intern-lab/Qwen-Image-2.1-rewriter-distill-trackio
License
This model is a derivative of Qwen-Image-2.1-PE-T2I outputs and is distributed under the Qwen Research License Agreement — see LICENSE (full text) and NOTICE (attribution per §3(c), modification statement per §3(b), and the Built with Qwen mark per §4(b)). Non-commercial use only.
- Downloads last month
- 4,350