GUI-Decisions-2B

Slot read vs writing JSON: the same model on the same Wikipedia task, 4 decisions in 0.63 s against 2.08 s

Slot reading with the 31B model (from the blog post): reading slots (left, 158 ms per decision) against writing the step as JSON (right, 520 ms). Try the 31B model in the Space.

mPLUG/GUI-Owl-1.5-2B-Instruct with the slot-reading LoRA merged in (bf16), which makes a computer-use step readable from one forward pass instead of decoded: the action, the point (256 bins per axis), a swipe's end, the key and the scroll, as next-token distributions at fixed slots. The small-model result of the blog post Stop Decoding Coordinates: 3× Faster Computer-Use Grounding. The LoRA alone: GUI-Decisions-2B-LoRA; the 31B version: GUI-Decisions-31B.

Results

Accuracy (blog post):

computer-use step (held-out AGUVIS, 1,600) ScreenSpot clicks (1,272)
GUI-Owl-2B + the LoRA 0.626 0.770
GUI-Owl-2B (its own tool-call format), native resolution 0.626 0.645
GUI-Owl-2B, same 1 MP images 0.610 0.626

Vanilla GUI-Owl's click numbers use its agent prompt, not its dedicated grounding prompt.

This merged model, read in-process with transformers (bf16): ScreenSpot clicks 0.763 against the published LoRA's 0.770 on the same cases (paired difference −0.006 [−0.013, +0.001], not significant; the unmerged LoRA read the same way scores 0.764), and computer-use steps 0.627 on the held-out AGUVIS set.

Latency of one computer-use step, measured with the LoRA on the base model (this merged model reads at the same speed: ~83 against ~87 ms end to end): vLLM 0.30, one RTX PRO 6000, cold screenshots, one request at a time, all conditions in one session (scripts/gui_owl.py bench 80):

how the step is decided median (p90)
slots, lazy: action, x, y, then only the slots the action needs (what decide does) 65 ms (71)
slots, the whole template (8 slots) 93 ms (104)
GUI-Owl-2B writing its tool call (~50 tokens), same 1 MP images 194 ms (217)
GUI-Owl-2B writing its tool call (~50 tokens), native resolution 323 ms (391)

The blog's earlier session measured 133, 225 and 444 ms for the last three rows on another box; the ratios agree (whole-template slots at ~30% of native-resolution GUI-Owl). End to end through decide, including the client's JPEG encoding, a step takes ~85 ms.

Use

With vLLM (recommended)

Served on one RTX PRO 6000 a computer-use step takes ~83 ms end to end through the code below (median over 60 cold screenshots, client JPEG encoding included; the LoRA on the base model: ~87 ms).

1. Install vLLM 0.30 and the reader (slotread, from the code repository):

git clone https://github.com/infinitylogesh/gui-decisions && cd gui-decisions
pip install "vllm==0.30.*" -e .
bash serving/patch_vllm.sh                 # recommended: 256 label ids per request

2. Serve:

vllm serve infinitylogesh/GUI-Decisions-2B --served-model-name gui-decisions-2b --port 8001 \
  --max-model-len 16384 --gpu-memory-utilization 0.6 --max-logprobs 300 --enable-prefix-caching \
  --mm-processor-cache-gb 0 --api-server-count 8 --limit-mm-per-prompt '{"image":1}'

(MODEL=infinitylogesh/GUI-Decisions-2B NAME=gui-decisions-2b MAX_LEN=16384 GPU_UTIL=0.6 EXTRA_ARGS='--limit-mm-per-prompt {"image":1}' bash serving/serve_vllm.sh runs the same command.)

3. Read steps:

from PIL import Image
from slotread.client import SlotModel

m = SlotModel("gui-decisions-2b", "infinitylogesh/GUI-Decisions-2B")   # server: VLLM_URL, default http://127.0.0.1:8001
step = m.decide(Image.open("screenshot.png"), "open the settings")         # {'action': 'click', 'x': ..., 'y': ..., 'ms': ...}

decide reads action, x and y first, then only the slots the action needs. Screenshots are downscaled to at most 1,048,576 pixels, as in training (image_cap in slotread_config.json; SlotModel applies it).

What a read is. For each slot, SlotModel sends one chat request whose assistant turn is prefilled up to that slot, and asks vLLM for the log-probabilities of exactly that slot's labels; nothing is generated beyond one token. The requests for one step run in parallel and share the image through vLLM's prefix cache:

{"model": "gui-decisions-2b",
 "messages": [{"role": "system", "content": [{"type": "text", "text": "<the task's system prompt, from slotread_config.json>"}]},
              {"role": "user", "content": [{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}},
                                           {"type": "text", "text": "Please generate the next move according to the UI screenshot, instruction and previous actions.\n\nInstruction: turn on dark mode\n\nPrevious actions:\nNone"}]},
              {"role": "assistant", "content": [{"type": "text", "text": "action: ___ x:"}]}],
 "add_generation_prompt": false, "continue_final_message": true, "chat_template_kwargs": {"enable_thinking": false},
 "max_tokens": 1, "temperature": 0, "logprobs": true, "top_logprobs": 1, "return_tokens_as_token_ids": true,
 "logprob_token_ids": ["<the 256 label ids of slot x, from slotread_config.json>"]}

The slot's answer is the softmax over those labels: the argmax for action; for a coordinate, the argmax bin refined by its ±2 neighbours (bin k of 256 is the value (k + 0.5) / 256 across the screen).

Notes:

  • patch_vllm.sh raises vLLM's per-request logprob_token_ids cap from 128 to 256, so a 256-bin slot is one request. Without it, set MAX_LABEL_IDS=128 for the client (each coordinate then takes two requests).
  • --api-server-count 8: a single API-server process serializes a step's parallel slot requests. --mm-processor-cache-gb 0: vLLM 0.30's media cache raced on concurrent requests for the same image.
  • previous takes the earlier steps as numbered lines, e.g. "1. click(One way)\n2. click(From)".

In-process with transformers (no server)

from PIL import Image
from slotread.local import LocalSlotModel

m = LocalSlotModel.from_pretrained("infinitylogesh/GUI-Decisions-2B")
print(m.decide(Image.open("screenshot.png"), "open the settings"))

Training

LoRA r 16, alpha 32, on the language model's attention and MLP projections; AGUVIS computer-use steps (5,000) and Wave-UI element clicks (2,000); one RTX PRO 6000, about 40 minutes.

Limits

One step at a time, no memory and no acting; computer use only (no boxes or masks). Text arguments (what to type, the final answer) need another writer: asked with our plain prompt, base GUI-Owl answered a city to type with the field's placeholder text. Trained on AGUVIS and Wave-UI; check their terms for your use.

Citation

@article{umapathi2026slotreading,
  title   = "Stop Decoding Coordinates: 3.5× Faster Computer-Use Grounding",
  author  = "Umapathi, Logesh Kumar",
  journal = "logeshumapathi.com",
  year    = "2026",
  month   = "Oct",
  url     = "https://logeshumapathi.com/blog/2026/10/06/slot-reading.html"
}
Downloads last month
37
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for infinitylogesh/GUI-Decisions-2B

Finetuned
(1)
this model
Quantizations
1 model

Collection including infinitylogesh/GUI-Decisions-2B

Article mentioning infinitylogesh/GUI-Decisions-2B