Darwin-27B-ZTC: A Single-Pass Judge and a Quantitative Look at Its Calibration
Two structural debts of LLM-as-judge
Using an LLM as a judge has quietly become standard: response grading, preference comparison, routing, safety filtering. The moment you use a generative LLM for this, two structural debts come with it.
The first is latency. Generation is O(L) sequential decoding in the output length L. Each judgment emits tens to hundreds of tokens, each conditioned on the previous one, so parallelism is limited. In a pipeline scoring millions of items, this cost accumulates linearly.
The second is calibration. What a judge actually needs is not just the verdict but how confident it is. To get confidence out of a generative model you either read logprobs (place the options in the prompt and read the token probabilities) or sample many times and count ratios. The former tends to diverge from the true posterior; the latter inflates variance and cost at once. Neither guarantees that the reported confidence matches the real probability of being correct.
Darwin-27B-ZTC (zero-token classifier) sidesteps both debts by reframing the problem. It treats judging not as generation but as conditional probability estimation, and returns a categorical distribution over each typed question in a single forward pass. Zero generated tokens, no decoding loop.
1. Problem formulation
Let the judge be a map f: (state s, question q) → Δ(Y_q), where Y_q is the answer space of question q and Δ(Y_q) is the probability simplex over it. The output is not a single label but a probability vector over possible answers.
A generative judge samples a token sequence t_1...t_L autoregressively and then back-derives a label and a confidence. ZTC emits an element of Δ(Y_q) directly. That one-line difference propagates into three properties:
- Determinism. No sampling step, so output variance is 0. The same input always returns the same distribution. Reported metrics are therefore exact values, not sample estimates, and need no confidence interval to reproduce.
- Computational cost. One forward pass per judgment. Latency does not scale with output length, so at batch scale this is a structural win, not just a constant-factor one.
- Confounder removal. Temperature, top-p, repetition penalty, max length all drop out of the evaluation. Cross-model comparison is fairer because setup differences no longer confound the numbers.
2. Evaluation setup: typed-decisions (zero-shot)
typed-decisions tests whether one judge can handle heterogeneous response types. There are three:
- noul: correctness of a free-form response.
- choice: selecting the right option among candidates.
- score: assigning a continuous or ordinal quality score.
The three have different decision boundaries. noul is open-ended correctness, choice is categorical classification, score is closer to ordinal regression. A single judgment representation covering all three is the point.
Measured pure zero-shot (the benchmark's train split was not used):
- Overall accuracy: 0.743 (general split), 2,000 judgments, 0 errors.
- By type: noul 0.847, choice 0.723, score 0.675.
- Calibration: KL 0.204, Brier 0.097.
The spread across types is itself informative. noul highest, score lowest matches the intuition that ordinal judgment is intrinsically harder than categorical judgment. That is why we report per-type rather than a single aggregate.
3. A quantitative look at calibration
Once a judge sits in front of a router or a safety gate, the evaluation criterion shifts from accuracy to calibration. The typical operation is: "if confidence is above threshold τ, auto-accept; otherwise defer to a larger model or a human." In this structure overconfidence is fatal. High confidence on a wrong answer leaks straight through the gate. Accuracy alone cannot evaluate a judge.
Brier score
For a predicted probability vector p and a one-hot outcome y, the mean squared error:
Brier = (1/N) Σ ‖ p_i − y_i ‖²
Brier is a strictly proper scoring rule: the unique strategy that optimizes expected score is to report the true probability. You cannot game it by inflating or deflating confidence. A value of 0.097 means predicted probabilities do not stray far from outcomes on average.
Brier also decomposes (Murphy):
Brier = reliability − resolution + uncertainty
reliability (lower is better) is the gap between predicted probability and observed frequency; resolution (higher is better) is the ability to separate cases into distinct risk groups. A low Brier is a compound signal: well calibrated and discriminative.
KL divergence
KL between the predicted distribution and the empirical outcome distribution:
KL(true ‖ pred) = Σ y log(y / p)
Closer to 0 means the two distributions are close. 0.204 indicates the prediction is near the outcome distribution.
Accuracy and calibration are independent axes
The key point is that accuracy and calibration are different axes. A model can be accurate but overconfident, or inaccurate but well calibrated. What matters operationally is achieving both at once. ZTC's 0.743 with (KL 0.204, Brier 0.097) reads as a signal on both the correctness axis and the confidence-honesty axis.
For a stricter view you want a reliability diagram (predicted probability vs observed accuracy per confidence bin) and ECE (Expected Calibration Error). Brier and KL are summary scalars; the reliability diagram shows which confidence bins are over- or under-confident, which is what you actually need when choosing τ.
4. A note on determinism and reproducibility
Every reported value is a deterministic output. It is not an average over samples, so it does not depend on seed or sample count and reproduces exactly on the same input. From a benchmark-integrity standpoint this matters: the thing being measured does not inject variance into its own output, so there is no "rerun and the score moves" problem. The standard-deviation term often reported for generative judges simply does not exist here.
5. Contrast with generative judging
| Axis | Generative judge | ZTC (single pass) |
|---|---|---|
| Getting confidence | logprob extraction or multi-sampling | distribution emitted directly |
| Latency | O(L) sequential decoding | one forward pass |
| Output variance | sampling variance present | 0 (deterministic) |
| Confounders | temperature, top-p, length | none |
| Calibration guarantee | indirect, uncertain | measured by a proper scoring rule |
6. Where it fits
- Large-scale output grading: at millions of items, a single forward pass per judgment is the decisive advantage.
- Model routing: branch easy inputs to small models and hard ones to large models by a confidence threshold τ. Better calibration makes τ more stable.
- Safety gating: the less overconfidence, the more aggressively you can set the auto-accept bar without leakage.
In all three, a deterministic, calibrated confidence is the center of the design, obtained without paying for generation.
7. Limits and open questions
- The numbers are on the general split of typed-decisions. Calibration under a very different input distribution needs separate validation, since calibration is sensitive to distribution shift.
- score accuracy is relatively low; losses or metrics that better reflect ordinal structure (e.g. rank correlation) are follow-up work.
- What is public: model weights, inference method, benchmark scores. Training configuration details are not disclosed.
Leaderboard
🏆 #1 on the official typed-decisions leaderboard on Hugging Face (0.743 accuracy, zero-shot).
- Model: https://huggingface.co/FINAL-Bench/Darwin-27B-ZTC
- Leaderboard: https://huggingface.co/datasets/LocalLLaMA/typed-decisions
Feedback is welcome on how to report reliability diagrams and ECE (what bin partition is appropriate), and on how to mitigate calibration degradation under distribution shift.

