Gender
Cross-model view of consistency for the gender attribute. Groups compared: man, woman, nonbinary. Evidence tiers compare each model against its own controls.
Consistency on gender, all models
Control band: the score range this metric produces under non-demographic substitutions (professions, arbitrary tags), the model's baseline substitution noise. How controls calibrate scores
Per-model detail
Leaderboard attribute cells deep-link to the rows below. Evidence tiers: how T0/T1/T2 are assigned.
| Model | Consistency | Tier | Worst-case gap | Control-adjusted | Refusal asym. (probe / incid.) |
|---|---|---|---|---|---|
| gemini-3.6-flash google/gemini-3.6-flash | 98.9 [97.0, 100.0] | T0 | 1.2 [0.0, 3.2] | 100.0 [99.4, 100.0] | 0.0 / 0.0 |
| mistral-medium-3-5 mistralai/mistral-medium-3-5 | 98.8 [96.6, 100.0] | T0 | 1.7 [0.0, 4.9] | 100.0 [100.0, 100.0] | 0.0 / 0.0 |
| gpt-5.6-sol openai/gpt-5.6-sol | 97.9 [94.7, 100.0] | T0 | 3.2 [0.0, 7.9] | 100.0 [100.0, 100.0] | 0.0 / 0.0 |
| gpt-5.6-terra openai/gpt-5.6-terra | 97.9 [95.6, 100.0] | T0 | 3.2 [0.0, 6.7] | 100.0 [100.0, 100.0] | 0.0 / 0.0 |
| claude-sonnet-5 anthropic/claude-sonnet-5 | 97.3 [93.7, 99.6] | T0 | 2.4 [0.5, 4.8] | 100.0 [96.6, 100.0] | 0.0 / 0.0 |
| gemini-3.1-pro-preview google/gemini-3.1-pro-preview | 96.4 [92.1, 100.0] | T0 | 5.4 [0.0, 11.9] | 99.7 [94.7, 100.0] | 0.0 / 0.0 |
| glm-5.2 z-ai/glm-5.2 | 95.7 [91.8, 98.9] | T0 | 4.2 [0.1, 9.8] | 100.0 [95.7, 100.0] | 0.0 / 0.0 |
| claude-opus-5 anthropic/claude-opus-5 | 95.7 [91.3, 99.2] | T0 | 4.9 [1.1, 10.1] | 100.0 [94.7, 100.0] | 0.0 / 0.0 |
| kimi-k3 moonshotai/kimi-k3 | 95.4 [92.6, 97.7] | T0 | 5.3 [2.2, 9.4] | 100.0 [98.5, 100.0] | 0.0 / 0.0 |
| grok-4.5 x-ai/grok-4.5 | 94.9 [90.7, 98.3] | T0 | 5.1 [2.0, 8.4] | 99.5 [94.5, 100.0] | 0.0 / 0.0 |
| qwen3.8-max qwen/qwen3.8-max | 94.1 [90.8, 97.1] | T0 | 8.1 [3.8, 12.9] | 100.0 [96.8, 100.0] | 0.0 / 0.0 |
| deepseek-v4-pro deepseek/deepseek-v4-pro | 93.9 [89.0, 98.1] | T0 | 7.8 [2.1, 15.0] | 100.0 [93.9, 100.0] | 0.0 / 0.0 |
| claude-fable-5 anthropic/claude-fable-5 | 92.5 [87.2, 97.3] | T0 | 9.6 [3.1, 17.4] | 94.9 [88.9, 100.0] | 0.0 / 2.6 |
| llama-4-maverick meta-llama/llama-4-maverick | 90.8 [83.3, 97.1] | T0 | 11.5 [3.7, 22.0] | 100.0 [94.4, 100.0] | 0.0 / 0.0 |
Directional lean: diagnostic only
Directional lean: diagnostic only. Shows which groups received more favorable responses in this sample. It is not a ranking, and it does not indicate intent or ideology.
Per-group lean by model: gender (diagnostic)
Diverging heatmap of per-group favorability lean for 14 models across 3 groups, centered at zero, range ±4.
| Model | man | woman | nonbinary |
|---|---|---|---|
| gemini-3.6-flash | −0.5 | +0.3 | +0.3 |
| mistral-medium-3-5 | −1.0 | +0.2 | +0.9 |
| gpt-5.6-sol | −0.8 | +1.6 | −0.8 |
| gpt-5.6-terra | −0.5 | −0.1 | +0.7 |
| claude-sonnet-5 | −0.8 | +0.2 | +0.6 |
| gemini-3.1-pro-preview | −0.2 | +2.0 | −1.6 |
| glm-5.2 | +0.2 | −1.2 | +1.0 |
| claude-opus-5 | −0.8 | +1.6 | −0.8 |
| kimi-k3 | +0.2 | +0.2 | −0.4 |
| grok-4.5 | +0.4 | +0.5 | −0.9 |
| qwen3.8-max | −1.9 | −0.4 | +2.3 |
| deepseek-v4-pro | −0.5 | +0.3 | +0.3 |
| claude-fable-5 | −2.5 | −1.4 | +3.9 |
| llama-4-maverick | −0.8 | −1.1 | +1.9 |