model accuracy
45 models · Standard
GPT-6 Astraxhigh
100.0% (89–100%)Claude Opus 5.5high
93.3% (79–98%)GPT-6 Solhigh
86.7% (70–95%)GPT-5.6 Solhigh
86.7% (70–95%)Gemini 3.8 Flashhigh
83.3% (66–93%)DeepSeek V4 Prohigh
83.3% (66–93%)Grok 4.7high
83.3% (66–93%)DeepSeek V4.1 Flashhigh
76.7% (59–88%)Muse Spark 1.3xhigh
76.7% (59–88%)Gemini 3.1 Pro Previewhigh
76.7% (59–88%)GPT-5.4xhigh
76.7% (59–88%)Kimi K3max
76.7% (59–88%)Claude Fable 5.1xhigh
76.7% (59–88%)Qwen3.8 Flashon
66.7% (49–81%)GPT-5.2high
66.7% (49–81%)Qwen3.8 Maxminimal
60.0% (42–75%)Claude Opus 4.5high
60.0% (42–75%)Gemini 3 Pro Previewhigh
56.7% (39–73%)Seed 2.1 Turboon
53.3% (36–70%)GPT-6 Lunamax
50.0% (33–67%)Gemini 3 Flash Previewhigh
50.0% (33–67%)DeepSeek V3.2 Specialehigh
46.7% (30–64%)MiMo V2.6 Proon
46.7% (30–64%)Qwen3.8 27Bxhigh
46.7% (30–64%)gpt-oss-120bhigh
43.3% (27–61%)DeepSeek V3.2high
43.3% (27–61%)Kimi K2.5high
43.3% (27–61%)GLM 5.3 Flashmax
40.0% (25–58%)GLM 5.3max
36.7% (22–54%)MiMo V2.6 Flashon
33.3% (19–51%)Seed 1.6high
33.3% (19–51%)Qwen3 Next 80B A3B Thinkingon
33.3% (19–51%)MiniMax M2.5high
33.3% (19–51%)Grok 4on
33.3% (19–51%)MiniMax M2.1high
30.0% (17–48%)GLM 5high
30.0% (17–48%)Claude Sonnet 4.5on
30.0% (17–48%)Mistral Medium 3.5low
30.0% (17–48%)Kimi K2 Thinkingon
26.7% (14–44%)Grok 4.1 Faston
23.3% (12–41%)MiMo V2 Flashhigh
20.0% (10–37%)GLM 4.7on
20.0% (10–37%)Seed 1.6 Flashhigh
13.3% (5–30%)OLMo 3.1 32B Thinkon
13.3% (5–30%)Kimi K2off
6.7% (2–21%)accuracy vs cost
Each point is a model variant. The stepped line traces the best measured tradeoffs.
Accuracy ↑↖ cheaper & more accurate
100%75%50%25%0%Seed 1.6 FlashDeepSeek V4.1 FlashGPT-6 Astra
$0.01$0.10
Cost per puzzle → (log scale: each solid line is 10× the last)
1 variant has no positive recorded cost value and cannot appear on a log scale.
View scatter data table
effort ladder
28 of 31 families scored higher at their top effort level than at their lowest. For 7 families, the top level was not their best.
GPT-6 Luna
offminimallowmediumhighxhighmax
GPT-6 Sol
offminimallowmediumhighxhighmax
Claude Opus 5.5
offminimallowmediumhighxhighmax
Grok 4.7
offminimallowmediumhighxhighmax
DeepSeek V4.1 Flash
offminimallowmediumhighxhighmax
GPT-6 Astra
offminimallowmediumhighxhighmax
Qwen3.8 Max
offminimallowmediumhighxhighmax
Muse Spark 1.3
offminimallowmediumhighxhighmax
Gemini 3.8 Flash
offminimallowmediumhighxhighmax
Claude Fable 5.1
offminimallowmediumhighxhighmax
GLM 5.3 Flash
offminimallowmediumhighxhighmax
GLM 5.3
offminimallowmediumhighxhighmax
Qwen3.8 27B
offminimallowmediumhighxhighmax
DeepSeek V4 Pro
offminimallowmediumhighxhighmax
Kimi K3
offminimallowmediumhighxhighmax
GPT-5.6 Sol
offminimallowmediumhighxhighmax
Mistral Medium 3.5
offminimallowmediumhighxhighmax
GPT-5.4
offminimallowmediumhighxhighmax
Gemini 3.1 Pro Preview
offminimallowmediumhighxhighmax
GLM 5
offminimallowmediumhighxhighmax
Kimi K2.5
offminimallowmediumhighxhighmax
GLM 4.7
offminimallowmediumhighxhighmax
Gemini 3 Flash Preview
offminimallowmediumhighxhighmax
GPT-5.2
offminimallowmediumhighxhighmax
DeepSeek V3.2
offminimallowmediumhighxhighmax
Claude Opus 4.5
offminimallowmediumhighxhighmax
gpt-oss-120b
offminimallowmediumhighxhighmax
DeepSeek V3.2 Speciale
offminimallowmediumhighxhighmax
Gemini 3 Pro Preview
offminimallowmediumhighxhighmax
Grok 4.1 Fast
offminimallowmediumhighxhighmax
MiMo V2 Flash
offminimallowmediumhighxhighmax
Reasoning off → on
Claude Sonnet 4.5
offon
View effort data table
detailed model statistics
Accuracy and usage for the selected tier.
| Model | 95% interval | 5×5 | 10×10 | 15×15 | Hard mode | ||
|---|---|---|---|---|---|---|---|
GPT-6 Astraxhigh | 100.0% (89–100%) | 100% | 100% | 100% | 50% | $10.82 | 1h 12m |
Claude Opus 5.5high | 93.3% (79–98%) | 100% | 100% | 80% | 80% | $5.24 | 48m 56s |
GPT-5.6 Solhigh | 86.7% (70–95%) | 100% | 100% | 60% | — | $2.27 | 49m 10s |
GPT-6 Solhigh | 86.7% (70–95%) | 100% | 90% | 70% | 0% | $1.86 | 52m 11s |
DeepSeek V4 Prohigh | 83.3% (66–93%) | 100% | 100% | 50% | 0% | $2.74 | 4h 20m |
Gemini 3.8 Flashhigh | 83.3% (66–93%) | 100% | 100% | 50% | 0% | $2.47 | 41m 54s |
Grok 4.7high | 83.3% (66–93%) | 100% | 90% | 60% | 0% | $8.97 | 7h 57m |
Claude Fable 5.1xhigh | 76.7% (59–88%) | 100% | 70% | 60% | 10% | $16.96 | 1h 01m |
DeepSeek V4.1 Flashhigh | 76.7% (59–88%) | 100% | 100% | 30% | 0% | $0.84 | 1h 23m |
Gemini 3.1 Pro Previewhigh | 76.7% (59–88%) | 100% | 90% | 40% | — | $5.43 | 1h 02m |
GPT-5.4xhigh | 76.7% (59–88%) | 100% | 90% | 40% | — | $10.20 | 2h 38m |
Kimi K3max | 76.7% (59–88%) | 100% | 70% | 60% | 0% | $14.52 | 6h 56m |
Muse Spark 1.3xhigh | 76.7% (59–88%) | 100% | 100% | 30% | 0% | $3.40 | 49m 53s |
GPT-5.2high | 66.7% (49–81%) | 100% | 80% | 20% | — | $10.86 | 4h 12m |
Qwen3.8 Flashon | 66.7% (49–81%) | 100% | 100% | 0% | 0% | $0.73 | 6h 16m |
Claude Opus 4.5high | 60.0% (42–75%) | 100% | 50% | 30% | — | $23.08 | 4h 45m |
Qwen3.8 Maxminimal | 60.0% (42–75%) | 100% | 70% | 10% | 0% | $4.82 | 5h 05m |
Gemini 3 Pro Previewhigh | 56.7% (39–73%) | 100% | 60% | 10% | — | $5.20 | 1h 20m |
Seed 2.1 Turboon | 53.3% (36–70%) | 100% | 50% | 10% | 0% | $3.13 | 4h 53m |
Gemini 3 Flash Previewhigh | 50.0% (33–67%) | 100% | 40% | 10% | — | $1.68 | 58m 45s |
GPT-6 Lunamax | 50.0% (33–67%) | 90% | 60% | 0% | 0% | $0.39 | 1h 50m |
DeepSeek V3.2 Specialehigh | 46.7% (30–64%) | 90% | 30% | 20% | — | $0.44 | 13h 02m |
MiMo V2.6 Proon | 46.7% (30–64%) | 80% | 60% | 0% | — | $1.00 | 6h 23m |
Qwen3.8 27Bxhigh | 46.7% (30–64%) | 100% | 40% | 0% | — | $3.61 | 7h 12m |
DeepSeek V3.2high | 43.3% (27–61%) | 100% | 20% | 10% | — | $0.62 | 5h 53m |
gpt-oss-120bhigh | 43.3% (27–61%) | 90% | 40% | 0% | — | $0.25 | 1h 13m |
Kimi K2.5high | 43.3% (27–61%) | 100% | 30% | 0% | — | $2.56 | 5h 25m |
GLM 5.3 Flashmax | 40.0% (25–58%) | 100% | 20% | 0% | — | $0.66 | 8h 29m |
GLM 5.3max | 36.7% (22–54%) | 80% | 20% | 10% | — | $4.72 | 4h 30m |
Grok 4on | 33.3% (19–51%) | 80% | 20% | 0% | — | $15.06 | 4h 49m |
MiMo V2.6 Flashon | 33.3% (19–51%) | 60% | 40% | 0% | — | $0.38 | 12h 11m |
MiniMax M2.5high | 33.3% (19–51%) | 90% | 10% | 0% | — | $1.49 | 7h 03m |
Qwen3 Next 80B A3B Thinkingon | 33.3% (19–51%) | 100% | 0% | 0% | — | $0.93 | 1h 14m |
Seed 1.6high | 33.3% (19–51%) | 100% | 0% | 0% | — | $0.83 | 2h 08m |
Claude Sonnet 4.5on | 30.0% (17–48%) | 90% | 0% | 0% | — | $5.98 | 1h 43m |
GLM 5high | 30.0% (17–48%) | 70% | 10% | 10% | — | $1.52 | 4h 07m |
MiniMax M2.1high | 30.0% (17–48%) | 90% | 0% | 0% | — | $0.61 | 2h 15m |
Mistral Medium 3.5low | 30.0% (17–48%) | 90% | 0% | 0% | — | $7.14 | 1h 20m |
Kimi K2 Thinkingon | 26.7% (14–44%) | 80% | 0% | 0% | — | $2.98 | 4h 31m |
Grok 4.1 Faston | 23.3% (12–41%) | 60% | 10% | 0% | — | $0.48 | 2h 02m |
GLM 4.7on | 20.0% (10–37%) | 60% | 0% | 0% | — | $1.13 | 3h 41m |
MiMo V2 Flashhigh | 20.0% (10–37%) | 60% | 0% | 0% | — | $0 | 3h 02m |
OLMo 3.1 32B Thinkon | 13.3% (5–30%) | 40% | 0% | 0% | — | $0.34 | 2h 04m |
Seed 1.6 Flashhigh | 13.3% (5–30%) | 40% | 0% | 0% | — | $0.20 | 1h 14m |
Kimi K2off | 6.7% (2–21%) | 20% | 0% | 0% | — | $0.23 | 53m 04s |
statistics by grid size
Combined results for the models shown.
5×545 models
88.0%
396/450 solved
1m 55s avg time$0.0265 avg cost
10×1045 models
45.8%
206/450 solved
9m 21s avg time$0.1590 avg cost
15×1545 models
19.3%
87/450 solved
10m 54s avg time$0.2340 avg cost
Hard mode (20×20)14 models
10.1%
14/139 solved
16m 05s avg time$0.7916 avg cost