Judges: default pair openai/gpt-5-mini + google/gemini-3.1-flash-lite; tiebreak anthropic/claude-haiku-4.5. A judge never scores a model from its own vendor family, so the effective panel is recorded per subject model. For 7 of the 14 models the rotation leaves no third judge, so a primary-pair disagreement there is recorded as contested straight away. Judging details
Judge agreement: Cohen's κ 0.57–1.00 by rubric; contested rate 3.4%. One label exceeds 90% prevalence on refusal_v1 (κ 0.57, PABAK 0.88, modal 93.6%); regard_v1 (κ 0.70, PABAK 0.95, modal 96.5%); κ is deflated there, so the gate reads PABAK (amended 2026-08-07). Contested rows are excluded from scores, never coerced.
Human spot check: 100 stratified items reviewed; 100.0% agreement with the judge panel (6 unsure, excluded from the rate). The sample was drawn before anthropic/claude-fable-5 was added to this run, so it covers the other models only. Protocol
Run full-2026-08 · bank sha256:8c9f5828… · harness 1c5654c7. Scores are comparable only within a bank hash. AIBiasBench measures consistency of treatment; it does not certify a model as unbiased.