The full statistical view: every score with its 95% confidence interval, per-attribute columns, control bands, and evidence tiers. The plain-language version of this run explains the same findings without the statistics.
Consistency of treatment across demographic groups
AIBiasBench asks each model matched pairs of prompts that differ only in a single group term and measures whether the answer changes. The headline metric is the consistency score (0–100; higher = more uniform treatment), read against each model's control band. This is a measurement report, not a scoreboard.
Control band: score range produced by non-demographic wording noise. No model can meaningfully exceed this band. How the band is measured
Score axis shows 80–100 of the 0–100 scale. Attribute cells are tinted by disparity magnitude (100 − consistency); deeper red = larger disparity. Click any column header to sort; click an attribute cell to compare all models on that attribute.
| Tier 1 | gpt-5.6-sol openai/gpt-5.6-sol #1 | 96.0 [94.7, 97.2]; control band 91.3–91.4 | 94.1 | 97.9 | 97.9 | 98.6 | 95.5 | 96.3 | 91.9 | 0.0% | 4,980 |
| Tier 1 | claude-opus-5 anthropic/claude-opus-5 #2 | 95.8 [94.3, 97.1]; control band 94.0–95.3 | 95.8 | 95.7 | 95.9 | 98.9 | 94.4 | 96.2 | 93.7 | 0.0% | 3,169 |
| Tier 1 | claude-fable-5 anthropic/claude-fable-5 #3 | 95.7 [93.8, 97.3]; control band 96.6–97.7 | 96.0 | 92.5 | 96.2 | 99.8 | 95.5 | 93.7 | 96.4 | 2.6% | 3,266 |
| Tier 1 | gemini-3.6-flash google/gemini-3.6-flash #4 | 94.9 [92.9, 96.7]; control band 95.6–96.0 | 96.1 | 98.9 | 98.4 | 95.5 | 93.3 | 90.6 | 91.3 | 0.0% | 4,347 |
| Tier 1 | claude-sonnet-5 anthropic/claude-sonnet-5 #5 | 94.7 [92.7, 96.5] †; control band 95.9–96.2 (inverted: For this model the arbitrary-tag control was noisier than the profession control, so the band's usual edges are swapped; it is drawn from the lower to the higher score.) | 93.2 | 97.3 | 99.6 | 97.0 | 94.6 | 89.0 | 92.2 | 0.0% | 4,383 |
| Tier 1 | gpt-5.6-terra openai/gpt-5.6-terra #6 | 94.3 [92.6, 95.9] †; control band 91.8–93.5 (inverted: For this model the arbitrary-tag control was noisier than the profession control, so the band's usual edges are swapped; it is drawn from the lower to the higher score.) | 93.5 | 97.9 | 95.2 | 96.3 | 93.3 | 93.3 | 90.7 | 0.0% | 4,936 |
| Tier 1 | grok-4.5 x-ai/grok-4.5 #7 | 94.2 [92.4, 95.8]; control band 95.2–95.4 | 95.7 | 94.9 | 94.6 | 94.4 | 95.4 | 92.9 | 91.1 | 0.0% | 5,795 |
| Tier 1 | gemini-3.1-pro-preview google/gemini-3.1-pro-preview #8 | 94.1 [92.0, 96.0]; control band 96.6–96.7 | 97.9 | 96.4 | 93.2 | 97.8 | 93.7 | 86.1 | 93.5 | 0.0% | 2,707 |
| Tier 1 | kimi-k3 moonshotai/kimi-k3 #9 | 93.9 [92.2, 95.4] †; control band 90.8–91.7 (inverted: For this model the arbitrary-tag control was noisier than the profession control, so the band's usual edges are swapped; it is drawn from the lower to the higher score.) | 93.1 | 95.4 | 94.4 | 96.0 | 95.8 | 91.8 | 90.8 | 0.9% | 5,663 |
| Tier 2 | mistral-medium-3-5 mistralai/mistral-medium-3-5 #10 | 93.0 [91.4, 94.6] †; control band 85.6–88.9 (inverted: For this model the arbitrary-tag control was noisier than the profession control, so the band's usual edges are swapped; it is drawn from the lower to the higher score.) | 94.0 | 98.8 | 95.6 | 89.4 | 94.2 | 91.2 | 88.1 | 1.3% | 2,723 |
| Tier 2 | glm-5.2 z-ai/glm-5.2 #11 | 92.6 [91.0, 94.2]; control band 93.5–95.1 | 92.7 | 95.7 | 95.3 | 92.0 | 92.7 | 88.9 | 91.2 | 0.0% | 5,713 |
| Tier 2 | qwen3.8-max qwen/qwen3.8-max #12 | 92.6 [90.8, 94.3]; control band 90.0–92.6 | 94.8 | 94.1 | 95.0 | 92.3 | 92.2 | 88.1 | 91.9 | 0.0% | 5,828 |
| Tier 2 | deepseek-v4-pro deepseek/deepseek-v4-pro #13 | 91.9 [90.1, 93.6] †; control band 92.5–93.3 (inverted: For this model the arbitrary-tag control was noisier than the profession control, so the band's usual edges are swapped; it is drawn from the lower to the higher score.) | 91.5 | 93.9 | 94.2 | 94.9 | 93.4 | 90.2 | 85.4 | 0.0% | 5,681 |
| Tier 2 | llama-4-maverick meta-llama/llama-4-maverick #14 | 89.9 [87.7, 92.1]; control band 82.9–85.5 | 91.9 | 90.8 | 91.6 | 90.1 | 92.2 | 87.0 | 85.7 | 0.9% | 5,094 |
Inverted control band (claude-sonnet-5, gpt-5.6-terra, kimi-k3, mistral-medium-3-5, deepseek-v4-pro): For this model the arbitrary-tag control was noisier than the profession control, so the band's usual edges are swapped; it is drawn from the lower to the higher score. How the band is measured
What this does not measure: real-world harm, model intent, or overall fairness. A high score means uniform treatment on this prompt bank, nothing more. Limitations
Metric definitions: consistency · control band · tiers · refusal asymmetry. Refusal asym. is the largest per-attribute gap between group refusal rates, in percentage points. n is the number of scored response rows, matching the banner at the top of the page.