Evaluating executive function from a single Wordle decision.
| Rank | Model | Overall Score | Win Rate | Avg Violations | API Error Rate | Penalty |
|---|---|---|---|---|---|---|
| 1 | gemini-3.1-pro-preview | 0.6650 | 8.0% | 0.000 | 0.0% | - |
| 2 | grok-4.20-reasoning | 0.6563 | 8.0% | 0.000 | 0.0% | - |
| 3 | claude-sonnet-4.6 | 0.6556 | 8.0% | 0.000 | 0.0% | - |
| 4 | gpt-5.5 | 0.6494 | 7.5% | 0.000 | 0.0% | - |
| 5 | gemma-4 | 0.6435 | 7.0% | 0.000 | 0.0% | - |
| 6 | glm-5 | 0.6347 | 7.0% | 0.000 | 0.0% | - |
| 7 | gemini-3-flash-preview | 0.6179 | 6.5% | 0.000 | 0.0% | - |
| 8 | qwen-3-next-80b-thinking | 0.6120 | 7.0% | 0.000 | 0.0% | - |
| 9 | gpt-5.4 | 0.5902 | 4.5% | 0.000 | 0.0% | - |
| 10 | claude-opus-4.6 | 0.5461 | 6.0% | 0.000 | 0.0% | - |
| 11 | gpt-5.4-mini | 0.5372 | 5.0% | 0.000 | 0.0% | - |
| 12 | qwen-3-next-80b-instruct | 0.5176 | 8.0% | 0.000 | 0.0% | - |
| 13 | deepseek-v3.2 | 0.5175 | 3.0% | 0.000 | 0.0% | - |
| 14 | grok-4.20-non-reasoning | 0.5058 | 2.0% | 0.000 | 0.0% | - |
| 15 | deepseek-r1 | 0.2836 | 5.5% | 0.000 | 0.0% | - |
Evaluating operational footprint. Sorted by a Combined Efficiency Score (33% Cost, 34% Speed, 33% Density).
| Rank | Model | Combined Score | Cost ($) | Bang-for-Buck | Time (Mins) | Velocity | Info Density |
|---|---|---|---|---|---|---|---|
| 1 | gpt-5.4-2026-03-05 | 68.5 / 100 | $1.0980 | 0.610 | 7.5m | 0.0888 | 0.295 |
| 2 | grok-4.20-0309-non-reasoning | 56.2 / 100 | $0.2292 | 2.225 | 5.4m | 0.0941 | 0.086 |
| 3 | gpt-5.4-mini-2026-03-17 | 54.4 / 100 | $0.3578 | 0.922 | 3.9m | 0.0855 | 0.163 |
| 4 | gemma-4-31b-it | 50.6 / 100 | $0.1610 | 5.839 | 184.7m | 0.0051 | 0.141 |
| 5 | gemini-3-flash-preview | 31.3 / 100 | $2.0203 | 0.386 | 52.7m | 0.0148 | 0.213 |
| 6 | gemini-3.1-pro-preview | 20.9 / 100 | $11.3193 | 0.081 | 136.2m | 0.0068 | 0.161 |
| 7 | gpt-5.5-2026-04-23 | 17.2 / 100 | $5.9839 | 0.167 | 45.7m | 0.0219 | 0.075 |
| 8 | deepseek-v3.2 | 14.6 / 100 | $0.5657 | 1.326 | 99.2m | 0.0076 | 0.039 |
| 9 | glm-5 | 12.5 / 100 | $1.1853 | 0.624 | 48.9m | 0.0151 | 0.031 |
| 10 | claude-opus-4-6-default | 10.8 / 100 | $8.4323 | 0.098 | 57.2m | 0.0145 | 0.045 |
| 11 | claude-sonnet-4-6-default | 9.6 / 100 | $6.0993 | 0.133 | 59.7m | 0.0136 | 0.036 |
| 12 | grok-4.20-0309-reasoning | 7.3 / 100 | $2.1145 | 0.312 | 51.5m | 0.0128 | 0.008 |
| 13 | qwen3-next-80b-a3b-thinking | 6.3 / 100 | $0.9209 | 0.586 | 92.0m | 0.0059 | 0.008 |
| 14 | gemma-3-4b-it | -25.1 / 100 | $0.0434 | -3.917 | 20.5m | -0.0083 | 0.000 |
| 15 | gemma-3-12b-it | -31.3 / 100 | $0.0350 | -5.143 | 29.0m | -0.0062 | 0.000 |