A same-vendor judge cannot grade its own family's work
Two judges from different model families agreed on 91% of items, and scored a hybrid question bank within 0.04 of the production one on a 3-point scale. The gap that decided it was citation quality, not correctness.
A blind, provenance-stripped benchmark of a hybrid Claude and GLM 5.2 question bank against a Claude-only production bank, scored independently by judges from two model families. An author and a judge from the same vendor share a blind spot, so a single same-family score tells you less than it appears to. Running the judges across families is what made the result defensible.
Macdara Ó Murchú - Built the Examworthy generation pipeline and its evaluation layer solo.
Published - the work it reports ran
Numbers fromcert-saas-factory/docs/eval-runs/saa-c03-live-vs-hybrid-benchmark-2026-06-23.mdcert-saas-factory/docs/eval-runs/glm-5.2-vs-deepseek-vs-claude-author-2026-06-22.md
- LLM-as-judge
- Evaluation
- Model selection
An eval that does not trust one vendor.
To validate a cheaper hybrid pipeline, I ran a blind, provenance-stripped benchmark on AWS SAA-C03: a hybrid Claude + GLM 5.2 bank against the Claude-only production bank, each scored independently by two judges from different model families. A same-vendor author and judge share a blind spot, so inter-judge agreement across families is the real signal.
| Claude: live | Claude: hybrid | GLM: live | GLM: hybrid | |
|---|---|---|---|---|
| Overall (of 3) | 2.99 | 2.95 | 2.94 | 2.92 |
| Clean rate | 96% | 90% | 98% | 94% |
| Correctness | 3.00 | 3.00 | 2.96 | 3.00 |
| References | 3.00 | 2.83 | 2.73 | 2.48 |
Blind scores, provenance stripped. 48 SAA-C03 questions per bank, 12 from each of four domains, judged by Claude Opus and GLM 5.2 on an identical 7-dimension rubric. Run 23 June 2026.
- 91% judge agreement
- The two judges, from different model families, agreed on 91% of items. Where independent judges converge, the verdict is robust.
- A dead heat
- Both judges scored the two banks within 0.02-0.04 on a 3-point scale. Correctness was effectively perfect in both - neither judge found a wrong answer in the hybrid bank.
- The honest differentiator
- The only material gap was citation quality, GLM's known weak spot. The benchmark surfaced it from the data rather than assuming it, and even caught a real pre-existing defect in the production bank.
- The decision
- A cheaper pipeline reaching production grade does not justify swapping a known-good bank for a larger one with weaker citations. Keep the bank, stay on the model that is ahead, and retest when the next genuinely capable lower-cost model lands.
Frontier versus the alternatives, decided with data.
Could a cheaper model author the questions? A separate assessment ran the candidates - local models on owned GPU hardware, plus cheaper cloud models - through the same deterministic gates and the same Opus audit as the frontier pipeline. The frontier model stayed ahead, so it stayed.
| Author model | Clean | Serious defects | Wrong keys |
|---|---|---|---|
| Claude Opus (frontier) | ~100% | 0% | 0 |
| GLM 5.2 (cloud) | 94.6% | 1.8% | 0 |
| DeepSeek V4 Pro (cloud) | 84.9% | 2.4% | 1 |
| Qwen3-14B (local, 16 GB GPU) | ~35% | ~13% | several |
Each model authored the same SY0-701 bank, then every question was independently audited by Claude Opus. Clean = no defects found; wrong keys = an incorrect option marked correct. Run 22 June 2026, a day before the blind benchmark above.
- Why the author role
- Generation is the largest share of the work, so the author is the one role where a cheaper model could actually move the numbers. That made it the first and best candidate to test.
- The hardware ceiling
- Owned GPU hardware caps how large a model runs at usable speed. Models up to 34B were tested, the larger ones only through slow offload - and even those that fit produced coherent, plausible-sounding wrong answers, the kind a weaker judge waves straight through.
- Correlated failure
- Pairing a local author with a same-family local judge is worse than it looks: they share the same blind spots, so the judge rubber-stamps the author's mistakes. Keeping the frontier model in the loop is what catches them.
What holds outside exam questions.
- Strip provenance before judging
- A judge told which pipeline produced a sample is scoring the label as much as the work. Stripping provenance costs nothing and is what makes the scores comparable at all.
- Cross the family line
- One judge produces a number. Two judges from different families produce a number you can defend, because their agreement rate is measurable and their blind spots are not shared.
- Expect the eval to find your own defects
- This run caught a pre-existing fault in the production bank it was meant to be the control for. An eval that only ever grades the challenger is not being pointed at enough of the system.
Have something to build?
Tell us what you are working on. We reply within a day, in plain language, with a clear next step.
Perth, Western Australia