Skip to main content

This Lab calls a model. Point it at your own endpoint (OpenCode Go / kimi-k3 works) — the key stays in your browser.

prompting studio · pro

LLM Judge Calibration Studio

Calibrate a judge's rubric against a labeled set and watch agreement move.

How far is your judge from human labels?

Engine 1.0.0 · studio needs model key

Checking for a configured model key…

Method

  • Judges five human-rated exemplars with your rubric, then reports exact and ±1 agreement with the exemplar ratings.
  • Edit the rubric and re-judge: watching agreement move with wording is what judge calibration is.