Age
Cross-model view of consistency for the age attribute. Groups compared: age 24, age 45, age 72. Evidence tiers compare each model against its own controls.
Consistency on age, all models
Control band: the score range this metric produces under non-demographic substitutions (professions, arbitrary tags), the model's baseline substitution noise. How controls calibrate scores
Per-model detail
Leaderboard attribute cells deep-link to the rows below. Evidence tiers: how T0/T1/T2 are assigned.
| Model | Consistency | Tier | Worst-case gap | Control-adjusted | Refusal asym. (probe / incid.) |
|---|---|---|---|---|---|
| claude-sonnet-5 anthropic/claude-sonnet-5 | 99.6 [99.1, 100.0] | T0 | 3.6 [0.0, 10.0] | 100.0 [100.0, 100.0] | 0.0 / 0.0 |
| gemini-3.6-flash google/gemini-3.6-flash | 98.4 [96.2, 100.0] | T0 | 1.9 [0.0, 4.7] | 100.0 [98.6, 100.0] | 0.0 / 0.0 |
| gpt-5.6-sol openai/gpt-5.6-sol | 97.9 [95.8, 99.5] | T0 | 2.4 [0.7, 4.4] | 100.0 [100.0, 100.0] | 0.0 / 0.0 |
| claude-fable-5 anthropic/claude-fable-5 | 96.2 [93.4, 98.6] | T0 | 5.2 [1.8, 9.1] | 98.5 [94.8, 100.0] | 0.0 / 0.0 |
| claude-opus-5 anthropic/claude-opus-5 | 95.9 [91.9, 99.3] | T0 | 4.5 [1.0, 9.0] | 100.0 [94.9, 100.0] | 0.0 / 0.0 |
| mistral-medium-3-5 mistralai/mistral-medium-3-5 | 95.6 [92.1, 98.6] | T0 | 6.3 [1.8, 12.1] | 100.0 [100.0, 100.0] | 0.0 / 0.0 |
| glm-5.2 z-ai/glm-5.2 | 95.3 [91.6, 98.2] | T0 | 3.8 [1.6, 6.0] | 100.0 [95.6, 100.0] | 0.0 / 0.0 |
| gpt-5.6-terra openai/gpt-5.6-terra | 95.2 [90.9, 98.6] | T0 | 3.2 [1.3, 5.3] | 100.0 [96.8, 100.0] | 0.0 / 0.0 |
| qwen3.8-max qwen/qwen3.8-max | 95.0 [91.2, 98.0] | T0 | 4.7 [1.6, 9.3] | 100.0 [97.5, 100.0] | 0.0 / 0.0 |
| grok-4.5 x-ai/grok-4.5 | 94.6 [89.8, 98.5] | T0 | 3.3 [1.2, 5.7] | 99.2 [93.3, 100.0] | 0.0 / 0.0 |
| kimi-k3 moonshotai/kimi-k3 | 94.4 [89.6, 98.0] | T0 | 5.6 [2.4, 9.9] | 100.0 [96.4, 100.0] | 0.0 / 0.0 |
| deepseek-v4-pro deepseek/deepseek-v4-pro | 94.2 [90.3, 97.2] | T0 | 6.6 [3.4, 10.3] | 100.0 [94.7, 100.0] | 0.0 / 0.0 |
| gemini-3.1-pro-preview google/gemini-3.1-pro-preview | 93.2 [85.3, 99.2] | T0 | 8.3 [1.2, 16.7] | 96.4 [87.5, 100.0] | 0.0 / 0.0 |
| llama-4-maverick meta-llama/llama-4-maverick | 91.6 [85.1, 97.2] | T0 | 6.2 [1.7, 11.6] | 100.0 [95.9, 100.0] | 0.0 / 0.9 |
Directional lean: diagnostic only
Directional lean: diagnostic only. Shows which groups received more favorable responses in this sample. It is not a ranking, and it does not indicate intent or ideology.
Per-group lean by model: age (diagnostic)
Diverging heatmap of per-group favorability lean for 14 models across 3 groups, centered at zero, range ±4.
| Model | age 24 | age 45 | age 72 |
|---|---|---|---|
| claude-sonnet-5 | −0.2 | −0.1 | +0.3 |
| gemini-3.6-flash | +0.1 | +0.1 | −0.1 |
| gpt-5.6-sol | −0.5 | −0.5 | +1.1 |
| claude-fable-5 | −1.2 | −1.0 | +2.1 |
| claude-opus-5 | −1.0 | −1.9 | +2.6 |
| mistral-medium-3-5 | +0.4 | −0.9 | +0.6 |
| glm-5.2 | +0.6 | −0.6 | +0.1 |
| gpt-5.6-terra | −0.7 | −0.9 | +1.7 |
| qwen3.8-max | −0.4 | −0.4 | +0.8 |
| grok-4.5 | −0.7 | +0.1 | +0.6 |
| kimi-k3 | −0.9 | +1.0 | −0.1 |
| deepseek-v4-pro | −0.4 | −0.5 | +0.9 |
| gemini-3.1-pro-preview | +0.6 | +0.9 | −1.5 |
| llama-4-maverick | −2.1 | +1.4 | +0.7 |