loading report…
Judge calibration — Cohen's κ
agreement vs human anchor set · chance-corrected
Pass-rate — Wilson 95% interval
gate reads the lower bound, never the point estimate
Judge vs human — confusion matrix
positive = "output passes"
Failure taxonomy — clustered
error-analysis-first · KMeans over embeddings
Judge drift — Cohen's κ over runs
κ re-scored against the frozen human anchor set on every run
judge κ
pass-rate
min κ threshold
drift (κ below threshold → CI blocked)
$ evalgate gate --min-pass-rate 0.90 --min-kappa 0.70