Orchidia V0 — Structured Decision Model on Qwen3.5-4B (Experimental Baseline)

Experimental V0 baseline. Not production-ready. Research artifact for Clef-inspired structured decisions.

Orchidia is a Clef-inspired structured decision model. It takes a STATE (text; image/video-ready) plus a DYNAMIC DECISION SCHEMA (arbitrary questions, types, and choices) and produces structured probability distributions for every question — not free-text generation.

  • Backbone: Qwen/Qwen3.5-4B (revision 851bf6e806ef at packaging).
  • V0 uses a FROZEN Qwen3.5-4B backbone. Only the Orchidia DynamicDecisionHead was trained. Full Qwen weights are NOT duplicated here; this repo contains the 10.5 MB head (2,628,609 params) plus configs/metrics/code to reconstruct V0 on top of the base model.
  • Dataset: ZefanCai/Open-Jev (release-v2-redistributable). V0 trained on the diverse/shuffled 2k experiment (see below).
  • Dynamic: any number of questions, any K per question. Output example: {"intent": {"refund": 0.43, ...}, "urgency": {...}, "escalate": {...}}.
  • Release status: experimental V0 baseline. The next step is balanced/diverse 8k on H100 (backbone initially frozen).

1. Model architecture

State + dynamic schema
→ Qwen3.5-4B trunk (Qwen3_5Model, hidden 2560, frozen)
→ question-conditioned representation (question text inside each input)
→ DynamicDecisionHead: LayerNorm(2560) + Linear(2560→1024) + GELU + Dropout(0.1) + Linear(1024→1)
→ per-option logits [sum(K)] (chunked forward, chunk=8)
→ per-question softmax → {qid: {choice: prob}}

There are NO fixed question classifiers. One shared scorer handles unseen questions and variable K. Loss: soft cross-entropy/KL for all kinds (choice/noul/score targets are distributions, including soft like [0,0.5,0.5]); ordinal EMD logged for score. Text-only V0 path; image/video hook preserved in code (explicit error until media batching is wired, no silent drop).

2. Training setup

V0 diverse/shuffled 2k (current baseline), frozen backbone, head-only:

Field Value
Base Qwen/Qwen3.5-4B rev 851bf6e806ef, frozen (requires_grad=False)
Data Open-Jev release-v2-redistributable, max_train 2000 shuffled (seed 0), max_val 256 shuffled (seed 1); diverse sources (painting 577, snake 331, security 228, vizdoom 164, … — not customer-only)
Steps 120 optimizer steps (executed 90 + 30 resumed after timeout; same data/seed, optimizer restarted on resume)
Batch / accum 2 questions/batch → 2×8=16 effective; forward_chunk_size 8
LR / optim / sched 1e-4 AdamW (head only), constant (none), clip_norm 1.0, GradScaler for fp16
Seq len 256 (truncate+pad per option-text; config default 512 for H100)
Precision fp16 on T4 (torch.amp); H100 → bf16 + flash_attention_2
Seed 0, deterministic cudnn
Checkpoints every 15 steps head-only + optimizer state; final checkpoints/v0_2k_shuf_final/head_final.pt + hf/ dir
Train val_loss 0.8638 (s30) → 0.8336 (s60) → 0.8042 (s90) → 0.8142 (final)

Full config: configs/v0_frozen.yaml. Checkpoint inspection: 6 tensors, fp32, shapes layer_norm.{weight,bias} [2560], mlp.0.weight [1024,2560], mlp.0.bias [1024], mlp.3.weight [1,1024], mlp.3.bias [1]; format {"head": state_dict, "config": {"hidden_size":2560,"backbone_id":"Qwen/Qwen3.5-4B"}}. See metrics/v0_2k_shuf.json.

3. Dataset information

ZefanCai/Open-Jev — typed decisions: kind ∈ {choice→categorical, noul→boolean, score→ordinal}, target distribution sums to 1.0. Rows grouped by group_id into {state, schema, labels}; splits are group-separated (verified 0 leakage train/val/test/ood). Counts: train 79116 / val 3723 / test 10356 / ood 15701. Do NOT redistribute Open-Jev data — this repo ships only Orchidia weights/configs/metrics. Use only state/question/kind/options as inputs (never target/metadata).

Source-bias warning (important): release-v2 is source-ordered. select(range(N)) without shuffle yields 100% customer-control-v1. V0 release uses deterministic shuffle(seed).select(N). Never use first-N.

4. Evaluation results

Meaningful = diverse/shuffled 2k (v0_2k_shuf_final/head_final.pt). Metrics: accuracy (argmax vs argmax incl. soft), macro/micro F1, ECE, Brier, per-kind/per-K, confidence. Evaluated n=512 per split.

Same first-512 slices (directly comparable):

Split Diverse (valid) loss/acc/F1m/ECE/Brier Biased (superseded, contrast only)
validation 0.756 / 0.703 / 0.714 / 0.119 / 0.423 0.974 / 0.551 / 0.458 / 0.235 / 0.570
test 0.499 / 0.873 / 0.869 / 0.175 / 0.248 0.223 / 0.930 / 0.929 / 0.044 / 0.093
ood 0.507 / 0.861 / 0.847 / 0.181 / 0.242 0.234 / 0.920 / 0.913 / 0.064 / 0.083

Hard diverse shuffled-512 (seed 1, K up to 16): val 0.806/0.631, test 0.750/0.688, ood 0.772/0.721.

Reading: biased test 0.93 is in-source memorization (train and test both customer-only; val K=5 acc 0.36). Diverse trades ~6pp in-source for +15pp validation and balanced cross-source — the honest baseline. Full tables in metrics/v0_2k_shuf.json. Eval code: src/evaluate.py (local) — release eval via inference/example.py + metrics file.

5. Known limitations

  • V0 head-only, 2k diverse (~120 steps): under-trained; expect gains from 8k.
  • Text-only V0; images/videos raise explicit error (hook exists, batching not wired).
  • max_length 256 truncates long states (vizdoom); use 512 on H100.
  • Calibration split unused; ECE/Brier reported only, no recalibration.
  • OOD in Open-Jev is source-specific (reserved wording/layouts/families), not universal domain shift.
  • T4 fp16 + reference attention kernels (no flash-linear-attention/causal_conv1d); slower than H100.
  • Resume loads head only (optimizer restarts) — fine for continuing, not exact resumption.

6. Exact base-model dependency

  • Qwen/Qwen3.5-4B, revision 851bf6e806ef (recorded at packaging; any compatible Qwen3.5-4B revision works but report yours).
  • transformers>=5.10 required (5.18.0 used; 5.0.0 lacks qwen3_5).
  • torch, datasets, huggingface_hub, peft (only for future LoRA), scikit-learn (metrics). See requirements.txt (local).
  • This repo does NOT contain Qwen weights. Reconstruction downloads Qwen from HF at inference time.

7. Installation instructions

pip install torch "transformers>=5.10" datasets huggingface_hub scikit-learn pyyaml tqdm
# Optional (H100 speed): pip install flash-attn  # then --attn flash_attention_2
# Optional (future LoRA): pip install peft accelerate
git clone https://huggingface.co/Seedyai/Orchidia && cd Orchidia
# Head downloads with the repo (checkpoints/v0_2k_shuf_final/head_final.pt).
# Qwen3.5-4B (~8GB) downloads on first inference via transformers.

No Open-Jev download needed for inference. For training/eval (local MiniClef-4B only): pip install -r requirements.txt and load_dataset("ZefanCai/Open-Jev", "release-v2-redistributable").

8. Inference example

Dynamic schema — any Q, any K (do not hard-code counts):

from inference.example import DynamicDecisionHead, predict_structured
# see inference/example.py for full runnable script
state = {"text": "Customer received the wrong product and wants a refund."}
schema = {"questions": [
  {"id": "intent", "type": "categorical", "question": "What is the customer intent?",
   "choices": ["refund", "shipping", "complaint", "other"]},
  {"id": "urgency", "type": "categorical", "question": "How urgent?",
   "choices": ["low", "medium", "high"]},
  {"id": "escalate", "type": "boolean", "question": "Should we escalate?",
   "choices": ["yes", "no"]},
]}
out = predict_structured(backbone, tokenizer, head, state, schema)
# {"intent": {"refund": 0.43, ...}, "urgency": {...}, "escalate": {...}}

Run: python inference/example.py --head checkpoints/v0_2k_shuf_final/head_final.pt --dtype fp16 (H100: --dtype bf16). Verified: Q=3 K=4/3/2 probs sum to 1.

9. Checkpoint loading example

import torch
from transformers import AutoProcessor, Qwen3_5Model
from inference.example import DynamicDecisionHead  # standalone copy; no local src needed

processor = AutoProcessor.from_pretrained("Qwen/Qwen3.5-4B", trust_remote_code=True)
tokenizer = processor.tokenizer
backbone = Qwen3_5Model.from_pretrained("Qwen/Qwen3.5-4B", dtype=torch.float16, trust_remote_code=True)
for p in backbone.parameters(): p.requires_grad = False  # V0 frozen
head = DynamicDecisionHead(hidden_size=2560)
ckpt = torch.load("checkpoints/v0_2k_shuf_final/head_final.pt", map_location="cpu")
head.load_state_dict(ckpt["head"])  # 2628609 params, 6 tensors, fp32
head.eval()

Reconstruction = Qwen trunk + this head + configs/v0_frozen.yaml hyperparams. Never train the backbone in V0.

10. Future roadmap

  1. Balanced/diverse 8k frozen training — DONE 2026-10-08 (T4, 500 steps, warm-started from 2k head; see §11).
  2. Evaluate test/OOD (first + shuffled, per-kind/per-K, calibration) — DONE for 8k (see §11).
  3. V0 16k diverse frozen — PREPARED FOR H100, NOT launched, do NOT run on T4 (configs/v0_16k_shuf.yaml: 512/bf16, warm-start from 8k head, exact-resume support in src/train.py) — NEXT.
  4. LoRA adaptation of Qwen3.5-4B (r=16, FEATURE_EXTRACTION task — proven 0.52% trainable).
  5. Compare frozen vs LoRA on identical data/seed/steps.
  6. Architecture/data improvements (longer context, stratified sampling, ordinal loss, media path, calibration use).

11. Current status (2026-10-08) — start here for continuation

PROJECT_STATUS.md is the canonical handoff. Read it first; this section is a pointer only.

  • Completed: 2k diverse (checkpoints/v0_2k_shuf_final/, NEVER overwrite) and 8k diverse (checkpoints/v0_8k_shuf/, 500 steps, FINAL val 0.7451; first-512 val/test/ood 0.738/0.875/0.881).
  • Next: launch 16k session 1 EXACTLY as specified in configs/v0_16k_shuf.yaml + PROJECT_STATUS.md "CURRENT NEXT STEP". Do NOT invent hyperparameters.
  • Resume rules (src/train.py): --resume head_final.pt = head-only warm start (step 0); --resume step-{N}.pt (+ sibling optim-{N}.pt) = exact resume. Never write into previous experiment dirs.

See PROJECT_STATUS.md (canonical handoff: experiments, decisions, bugs, workflow). Local dev repo MiniClef-4B retains src/, scripts/, full logs; this HF repo is the minimal reconstructable release.


License/attribution: Orchidia code/configs/metrics as released here; Qwen3.5-4B weights follow their upstream license; Open-Jev data NOT included (see its repo for license; CC0 synthetic + third-party notes). This V0 is experimental.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Seedyai/Orchidia

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(917)
this model