Religion
Cross-model view of consistency for the religion attribute. Groups compared: christian, muslim, jewish, hindu, buddhist, atheist. Evidence tiers compare each model against its own controls.
Consistency on religion, all models
Control band: the score range this metric produces under non-demographic substitutions (professions, arbitrary tags), the model's baseline substitution noise. How controls calibrate scores
Per-model detail
Leaderboard attribute cells deep-link to the rows below. Evidence tiers: how T0/T1/T2 are assigned.
| Model | Consistency | Tier | Worst-case gap | Control-adjusted | Refusal asym. (probe / incid.) |
|---|---|---|---|---|---|
| claude-fable-5 anthropic/claude-fable-5 | 99.8 [99.4, 100.0] | T0 | 0.4 [0.0, 1.2] | 100.0 [99.8, 100.0] | 0.0 / 1.3 |
| claude-opus-5 anthropic/claude-opus-5 | 98.9 [97.7, 99.9] | T0 | 3.2 [0.1, 7.9] | 100.0 [99.8, 100.0] | 0.0 / 0.0 |
| gpt-5.6-sol openai/gpt-5.6-sol | 98.6 [97.0, 99.8] | T0 | 3.8 [0.6, 7.6] | 100.0 [100.0, 100.0] | 0.0 / 0.0 |
| gemini-3.1-pro-preview google/gemini-3.1-pro-preview | 97.8 [95.1, 100.0] | T0 | 4.8 [0.0, 10.6] | 100.0 [97.1, 100.0] | 0.0 / 0.0 |
| claude-sonnet-5 anthropic/claude-sonnet-5 | 97.0 [93.6, 99.4] | T0 | 11.6 [1.9, 24.7] | 100.0 [96.4, 100.0] | 0.0 / 0.0 |
| gpt-5.6-terra openai/gpt-5.6-terra | 96.3 [93.1, 99.0] | T0 | 9.7 [2.6, 20.4] | 100.0 [98.5, 100.0] | 0.0 / 0.0 |
| kimi-k3 moonshotai/kimi-k3 | 96.0 [92.5, 98.7] | T0 | 7.0 [2.2, 13.0] | 100.0 [98.7, 100.0] | 0.0 / 0.0 |
| gemini-3.6-flash google/gemini-3.6-flash | 95.5 [91.6, 98.8] | T0 | 7.5 [2.4, 12.6] | 99.5 [94.6, 100.0] | 0.0 / 0.0 |
| deepseek-v4-pro deepseek/deepseek-v4-pro | 94.9 [91.8, 97.7] | T0 | 8.2 [3.7, 13.6] | 100.0 [95.9, 100.0] | 0.0 / 0.0 |
| grok-4.5 x-ai/grok-4.5 | 94.4 [90.2, 98.1] | T0 | 9.0 [2.7, 15.6] | 99.0 [94.0, 100.0] | 0.0 / 0.0 |
| qwen3.8-max qwen/qwen3.8-max | 92.3 [87.7, 96.6] | T0 | 13.9 [6.1, 22.4] | 99.7 [94.0, 100.0] | 0.0 / 0.0 |
| glm-5.2 z-ai/glm-5.2 | 92.0 [88.1, 95.5] | T0 | 14.3 [7.8, 22.0] | 96.9 [92.3, 100.0] | 0.0 / 0.0 |
| llama-4-maverick meta-llama/llama-4-maverick | 90.1 [85.7, 94.3] | T0 | 19.5 [11.5, 28.7] | 100.0 [95.3, 100.0] | 0.0 / 0.0 |
| mistral-medium-3-5 mistralai/mistral-medium-3-5 | 89.4 [85.2, 93.5] | T0 | 25.3 [15.0, 36.0] | 100.0 [94.6, 100.0] | 0.0 / 1.3 |
Directional lean: diagnostic only
Directional lean: diagnostic only. Shows which groups received more favorable responses in this sample. It is not a ranking, and it does not indicate intent or ideology.
Per-group lean by model: religion (diagnostic)
Diverging heatmap of per-group favorability lean for 14 models across 6 groups, centered at zero, range ±8.
| Model | christian | muslim | jewish | hindu | buddhist | atheist |
|---|---|---|---|---|---|---|
| claude-fable-5 | −0.2 | +0.1 | −0.1 | +0.2 | −0.1 | 0.0 |
| claude-opus-5 | −0.4 | +0.8 | −0.4 | +0.3 | +0.7 | −1.0 |
| gpt-5.6-sol | −0.3 | −1.1 | −0.7 | +0.5 | +0.6 | +1.1 |
| gemini-3.1-pro-preview | −1.8 | +0.4 | −0.4 | +0.3 | −0.7 | +2.6 |
| claude-sonnet-5 | +0.7 | +0.2 | −0.4 | 0.0 | −1.1 | +0.7 |
| gpt-5.6-terra | +0.4 | −0.9 | +1.4 | −1.6 | −0.2 | +1.0 |
| kimi-k3 | −0.6 | +1.9 | −0.2 | +0.1 | −1.9 | +0.6 |
| gemini-3.6-flash | −2.2 | −2.4 | +1.0 | +0.3 | +0.6 | +2.6 |
| deepseek-v4-pro | −1.1 | −1.1 | +1.1 | −1.3 | +1.4 | +0.9 |
| grok-4.5 | −1.2 | +0.9 | +1.2 | −0.3 | +0.1 | −0.6 |
| qwen3.8-max | −0.4 | −1.7 | −3.6 | +1.0 | +2.3 | +2.3 |
| glm-5.2 | −1.6 | +1.0 | −0.8 | −1.7 | −1.0 | +4.0 |
| llama-4-maverick | −1.3 | +1.9 | −2.4 | +0.5 | +1.2 | −0.1 |
| mistral-medium-3-5 | −3.3 | −3.8 | −2.3 | +0.5 | +1.7 | +7.0 |