Skip to content

Bank v1.0.0 · evaluated 2026-08-07 · n = 2,707–5,828 scored rows per model · sha256:8c9f5828

Age

Cross-model view of consistency for the age attribute. Groups compared: age 24, age 45, age 72. Evidence tiers compare each model against its own controls.

Consistency on age, all models

The shaded band is the run-average control band; each model's own band is on its model page. Ordering within overlapping CIs is not statistically meaningful. Tiers

Control band: the score range this metric produces under non-demographic substitutions (professions, arbitrary tags), the model's baseline substitution noise. How controls calibrate scores

Per-model detail

Leaderboard attribute cells deep-link to the rows below. Evidence tiers: how T0/T1/T2 are assigned.

ModelConsistencyTierWorst-case gapControl-adjustedRefusal asym. (probe / incid.)
claude-sonnet-5 anthropic/claude-sonnet-599.6 [99.1, 100.0]T03.6 [0.0, 10.0]100.0 [100.0, 100.0]0.0 / 0.0
gemini-3.6-flash google/gemini-3.6-flash98.4 [96.2, 100.0]T01.9 [0.0, 4.7]100.0 [98.6, 100.0]0.0 / 0.0
gpt-5.6-sol openai/gpt-5.6-sol97.9 [95.8, 99.5]T02.4 [0.7, 4.4]100.0 [100.0, 100.0]0.0 / 0.0
claude-fable-5 anthropic/claude-fable-596.2 [93.4, 98.6]T05.2 [1.8, 9.1]98.5 [94.8, 100.0]0.0 / 0.0
claude-opus-5 anthropic/claude-opus-595.9 [91.9, 99.3]T04.5 [1.0, 9.0]100.0 [94.9, 100.0]0.0 / 0.0
mistral-medium-3-5 mistralai/mistral-medium-3-595.6 [92.1, 98.6]T06.3 [1.8, 12.1]100.0 [100.0, 100.0]0.0 / 0.0
glm-5.2 z-ai/glm-5.295.3 [91.6, 98.2]T03.8 [1.6, 6.0]100.0 [95.6, 100.0]0.0 / 0.0
gpt-5.6-terra openai/gpt-5.6-terra95.2 [90.9, 98.6]T03.2 [1.3, 5.3]100.0 [96.8, 100.0]0.0 / 0.0
qwen3.8-max qwen/qwen3.8-max95.0 [91.2, 98.0]T04.7 [1.6, 9.3]100.0 [97.5, 100.0]0.0 / 0.0
grok-4.5 x-ai/grok-4.594.6 [89.8, 98.5]T03.3 [1.2, 5.7]99.2 [93.3, 100.0]0.0 / 0.0
kimi-k3 moonshotai/kimi-k394.4 [89.6, 98.0]T05.6 [2.4, 9.9]100.0 [96.4, 100.0]0.0 / 0.0
deepseek-v4-pro deepseek/deepseek-v4-pro94.2 [90.3, 97.2]T06.6 [3.4, 10.3]100.0 [94.7, 100.0]0.0 / 0.0
gemini-3.1-pro-preview google/gemini-3.1-pro-preview93.2 [85.3, 99.2]T08.3 [1.2, 16.7]96.4 [87.5, 100.0]0.0 / 0.0
llama-4-maverick meta-llama/llama-4-maverick91.6 [85.1, 97.2]T06.2 [1.7, 11.6]100.0 [95.9, 100.0]0.0 / 0.9

Directional lean: diagnostic only

Directional lean: diagnostic only. Shows which groups received more favorable responses in this sample. It is not a ranking, and it does not indicate intent or ideology.

Per-group lean by model: age (diagnostic)

Diverging heatmap of per-group favorability lean for 14 models across 3 groups, centered at zero, range ±4.

Modelage 24age 45age 72
claude-sonnet-5−0.2−0.1+0.3
gemini-3.6-flash+0.1+0.1−0.1
gpt-5.6-sol−0.5−0.5+1.1
claude-fable-5−1.2−1.0+2.1
claude-opus-5−1.0−1.9+2.6
mistral-medium-3-5+0.4−0.9+0.6
glm-5.2+0.6−0.6+0.1
gpt-5.6-terra−0.7−0.9+1.7
qwen3.8-max−0.4−0.4+0.8
grok-4.5−0.7+0.1+0.6
kimi-k3−0.9+1.0−0.1
deepseek-v4-pro−0.4−0.5+0.9
gemini-3.1-pro-preview+0.6+0.9−1.5
llama-4-maverick−2.1+1.4+0.7
Lean:≤ −2−2…−0±0+0…+2≥ +2(purple = below group mean, teal = above; neither is a verdict)
Rows follow this attribute's consistency order, the same order as the dot plot and the table above, and cannot be sorted. Purple: below the all-group mean; teal: above. Neither is a verdict. Definition

Back to leaderboard · All attributes