Skip to content

Bank v1.0.0 · evaluated 2026-08-07 · n = 2,707–5,828 scored rows per model · sha256:8c9f5828

Religion

Cross-model view of consistency for the religion attribute. Groups compared: christian, muslim, jewish, hindu, buddhist, atheist. Evidence tiers compare each model against its own controls.

Consistency on religion, all models

The shaded band is the run-average control band; each model's own band is on its model page. Ordering within overlapping CIs is not statistically meaningful. Tiers

Control band: the score range this metric produces under non-demographic substitutions (professions, arbitrary tags), the model's baseline substitution noise. How controls calibrate scores

Per-model detail

Leaderboard attribute cells deep-link to the rows below. Evidence tiers: how T0/T1/T2 are assigned.

ModelConsistencyTierWorst-case gapControl-adjustedRefusal asym. (probe / incid.)
claude-fable-5 anthropic/claude-fable-599.8 [99.4, 100.0]T00.4 [0.0, 1.2]100.0 [99.8, 100.0]0.0 / 1.3
claude-opus-5 anthropic/claude-opus-598.9 [97.7, 99.9]T03.2 [0.1, 7.9]100.0 [99.8, 100.0]0.0 / 0.0
gpt-5.6-sol openai/gpt-5.6-sol98.6 [97.0, 99.8]T03.8 [0.6, 7.6]100.0 [100.0, 100.0]0.0 / 0.0
gemini-3.1-pro-preview google/gemini-3.1-pro-preview97.8 [95.1, 100.0]T04.8 [0.0, 10.6]100.0 [97.1, 100.0]0.0 / 0.0
claude-sonnet-5 anthropic/claude-sonnet-597.0 [93.6, 99.4]T011.6 [1.9, 24.7]100.0 [96.4, 100.0]0.0 / 0.0
gpt-5.6-terra openai/gpt-5.6-terra96.3 [93.1, 99.0]T09.7 [2.6, 20.4]100.0 [98.5, 100.0]0.0 / 0.0
kimi-k3 moonshotai/kimi-k396.0 [92.5, 98.7]T07.0 [2.2, 13.0]100.0 [98.7, 100.0]0.0 / 0.0
gemini-3.6-flash google/gemini-3.6-flash95.5 [91.6, 98.8]T07.5 [2.4, 12.6]99.5 [94.6, 100.0]0.0 / 0.0
deepseek-v4-pro deepseek/deepseek-v4-pro94.9 [91.8, 97.7]T08.2 [3.7, 13.6]100.0 [95.9, 100.0]0.0 / 0.0
grok-4.5 x-ai/grok-4.594.4 [90.2, 98.1]T09.0 [2.7, 15.6]99.0 [94.0, 100.0]0.0 / 0.0
qwen3.8-max qwen/qwen3.8-max92.3 [87.7, 96.6]T013.9 [6.1, 22.4]99.7 [94.0, 100.0]0.0 / 0.0
glm-5.2 z-ai/glm-5.292.0 [88.1, 95.5]T014.3 [7.8, 22.0]96.9 [92.3, 100.0]0.0 / 0.0
llama-4-maverick meta-llama/llama-4-maverick90.1 [85.7, 94.3]T019.5 [11.5, 28.7]100.0 [95.3, 100.0]0.0 / 0.0
mistral-medium-3-5 mistralai/mistral-medium-3-589.4 [85.2, 93.5]T025.3 [15.0, 36.0]100.0 [94.6, 100.0]0.0 / 1.3

Directional lean: diagnostic only

Directional lean: diagnostic only. Shows which groups received more favorable responses in this sample. It is not a ranking, and it does not indicate intent or ideology.

Per-group lean by model: religion (diagnostic)

Diverging heatmap of per-group favorability lean for 14 models across 6 groups, centered at zero, range ±8.

Modelchristianmuslimjewishhindubuddhistatheist
claude-fable-5−0.2+0.1−0.1+0.2−0.10.0
claude-opus-5−0.4+0.8−0.4+0.3+0.7−1.0
gpt-5.6-sol−0.3−1.1−0.7+0.5+0.6+1.1
gemini-3.1-pro-preview−1.8+0.4−0.4+0.3−0.7+2.6
claude-sonnet-5+0.7+0.2−0.40.0−1.1+0.7
gpt-5.6-terra+0.4−0.9+1.4−1.6−0.2+1.0
kimi-k3−0.6+1.9−0.2+0.1−1.9+0.6
gemini-3.6-flash−2.2−2.4+1.0+0.3+0.6+2.6
deepseek-v4-pro−1.1−1.1+1.1−1.3+1.4+0.9
grok-4.5−1.2+0.9+1.2−0.3+0.1−0.6
qwen3.8-max−0.4−1.7−3.6+1.0+2.3+2.3
glm-5.2−1.6+1.0−0.8−1.7−1.0+4.0
llama-4-maverick−1.3+1.9−2.4+0.5+1.2−0.1
mistral-medium-3-5−3.3−3.8−2.3+0.5+1.7+7.0
Lean:≤ −4−4…−1±1+1…+4≥ +4(purple = below group mean, teal = above; neither is a verdict)
Rows follow this attribute's consistency order, the same order as the dot plot and the table above, and cannot be sorted. Purple: below the all-group mean; teal: above. Neither is a verdict. Definition

Back to leaderboard · All attributes