Skip to content

Bank v1.0.0 · evaluated 2026-08-07 · n = 2,707–5,828 scored rows per model · sha256:8c9f5828

The full statistical view: every score with its 95% confidence interval, per-attribute columns, control bands, and evidence tiers. The plain-language version of this run explains the same findings without the statistics.

Consistency of treatment across demographic groups

AIBiasBench asks each model matched pairs of prompts that differ only in a single group term and measures whether the answer changes. The headline metric is the consistency score (0–100; higher = more uniform treatment), read against each model's control band. This is a measurement report, not a scoreboard.

Control band: score range produced by non-demographic wording noise. No model can meaningfully exceed this band. How the band is measured

Score axis shows 80–100 of the 0–100 scale. Attribute cells are tinted by disparity magnitude (100 − consistency); deeper red = larger disparity. Click any column header to sort; click an attribute cell to compare all models on that attribute.

Tier 1gpt-5.6-sol openai/gpt-5.6-sol #1
96.0 [94.7, 97.2]; control band 91.3–91.4
94.197.997.998.695.596.391.90.0%4,980
Tier 1claude-opus-5 anthropic/claude-opus-5 #2
95.8 [94.3, 97.1]; control band 94.0–95.3
95.895.795.998.994.496.293.70.0%3,169
Tier 1claude-fable-5 anthropic/claude-fable-5 #3
95.7 [93.8, 97.3]; control band 96.6–97.7
96.092.596.299.895.593.796.42.6%3,266
Tier 1gemini-3.6-flash google/gemini-3.6-flash #4
94.9 [92.9, 96.7]; control band 95.6–96.0
96.198.998.495.593.390.691.30.0%4,347
Tier 1claude-sonnet-5 anthropic/claude-sonnet-5 #5
94.7 [92.7, 96.5] ; control band 95.9–96.2 (inverted: For this model the arbitrary-tag control was noisier than the profession control, so the band's usual edges are swapped; it is drawn from the lower to the higher score.)
93.297.399.697.094.689.092.20.0%4,383
Tier 1gpt-5.6-terra openai/gpt-5.6-terra #6
94.3 [92.6, 95.9] ; control band 91.8–93.5 (inverted: For this model the arbitrary-tag control was noisier than the profession control, so the band's usual edges are swapped; it is drawn from the lower to the higher score.)
93.597.995.296.393.393.390.70.0%4,936
Tier 1grok-4.5 x-ai/grok-4.5 #7
94.2 [92.4, 95.8]; control band 95.2–95.4
95.794.994.694.495.492.991.10.0%5,795
Tier 1gemini-3.1-pro-preview google/gemini-3.1-pro-preview #8
94.1 [92.0, 96.0]; control band 96.6–96.7
97.996.493.297.893.786.193.50.0%2,707
Tier 1kimi-k3 moonshotai/kimi-k3 #9
93.9 [92.2, 95.4] ; control band 90.8–91.7 (inverted: For this model the arbitrary-tag control was noisier than the profession control, so the band's usual edges are swapped; it is drawn from the lower to the higher score.)
93.195.494.496.095.891.890.80.9%5,663
Tier 2mistral-medium-3-5 mistralai/mistral-medium-3-5 #10
93.0 [91.4, 94.6] ; control band 85.6–88.9 (inverted: For this model the arbitrary-tag control was noisier than the profession control, so the band's usual edges are swapped; it is drawn from the lower to the higher score.)
94.098.895.689.494.291.288.11.3%2,723
Tier 2glm-5.2 z-ai/glm-5.2 #11
92.6 [91.0, 94.2]; control band 93.5–95.1
92.795.795.392.092.788.991.20.0%5,713
Tier 2qwen3.8-max qwen/qwen3.8-max #12
92.6 [90.8, 94.3]; control band 90.0–92.6
94.894.195.092.392.288.191.90.0%5,828
Tier 2deepseek-v4-pro deepseek/deepseek-v4-pro #13
91.9 [90.1, 93.6] ; control band 92.5–93.3 (inverted: For this model the arbitrary-tag control was noisier than the profession control, so the band's usual edges are swapped; it is drawn from the lower to the higher score.)
91.593.994.294.993.490.285.40.0%5,681
Tier 2llama-4-maverick meta-llama/llama-4-maverick #14
89.9 [87.7, 92.1]; control band 82.9–85.5
91.990.891.690.192.287.085.70.9%5,094

Inverted control band (claude-sonnet-5, gpt-5.6-terra, kimi-k3, mistral-medium-3-5, deepseek-v4-pro): For this model the arbitrary-tag control was noisier than the profession control, so the band's usual edges are swapped; it is drawn from the lower to the higher score. How the band is measured

What this does not measure: real-world harm, model intent, or overall fairness. A high score means uniform treatment on this prompt bank, nothing more. Limitations

Metric definitions: consistency · control band · tiers · refusal asymmetry. Refusal asym. is the largest per-attribute gap between group refusal rates, in percentage points. n is the number of scored response rows, matching the banner at the top of the page.