nonobench results

v1.2

How well LLMs solve nonogram puzzles. Compare accuracy, speed and cost across grid sizes.

45 models · 110 variants with solved resultsUpdated 9/27/2026

model accuracy

45 models · Standard
GPT-6 Astraxhigh
100.0% (89–100%)
Claude Opus 5.5high
93.3% (79–98%)
GPT-6 Solhigh
86.7% (70–95%)
GPT-5.6 Solhigh
86.7% (70–95%)
Gemini 3.8 Flashhigh
83.3% (66–93%)
DeepSeek V4 Prohigh
83.3% (66–93%)
Grok 4.7high
83.3% (66–93%)
DeepSeek V4.1 Flashhigh
76.7% (59–88%)
Muse Spark 1.3xhigh
76.7% (59–88%)
Gemini 3.1 Pro Previewhigh
76.7% (59–88%)
GPT-5.4xhigh
76.7% (59–88%)
Kimi K3max
76.7% (59–88%)
Claude Fable 5.1xhigh
76.7% (59–88%)
Qwen3.8 Flashon
66.7% (49–81%)
GPT-5.2high
66.7% (49–81%)
Qwen3.8 Maxminimal
60.0% (42–75%)
Claude Opus 4.5high
60.0% (42–75%)
Gemini 3 Pro Previewhigh
56.7% (39–73%)
Seed 2.1 Turboon
53.3% (36–70%)
GPT-6 Lunamax
50.0% (33–67%)
Gemini 3 Flash Previewhigh
50.0% (33–67%)
DeepSeek V3.2 Specialehigh
46.7% (30–64%)
MiMo V2.6 Proon
46.7% (30–64%)
Qwen3.8 27Bxhigh
46.7% (30–64%)
gpt-oss-120bhigh
43.3% (27–61%)
DeepSeek V3.2high
43.3% (27–61%)
Kimi K2.5high
43.3% (27–61%)
GLM 5.3 Flashmax
40.0% (25–58%)
GLM 5.3max
36.7% (22–54%)
MiMo V2.6 Flashon
33.3% (19–51%)
Seed 1.6high
33.3% (19–51%)
Qwen3 Next 80B A3B Thinkingon
33.3% (19–51%)
MiniMax M2.5high
33.3% (19–51%)
Grok 4on
33.3% (19–51%)
MiniMax M2.1high
30.0% (17–48%)
GLM 5high
30.0% (17–48%)
Claude Sonnet 4.5on
30.0% (17–48%)
Mistral Medium 3.5low
30.0% (17–48%)
Kimi K2 Thinkingon
26.7% (14–44%)
Grok 4.1 Faston
23.3% (12–41%)
MiMo V2 Flashhigh
20.0% (10–37%)
GLM 4.7on
20.0% (10–37%)
Seed 1.6 Flashhigh
13.3% (5–30%)
OLMo 3.1 32B Thinkon
13.3% (5–30%)
Kimi K2off
6.7% (2–21%)

accuracy vs cost

Each point is a model variant. The stepped line traces the best measured tradeoffs.

Accuracy ↑↖ cheaper & more accurate
100%75%50%25%0%
Seed 1.6 FlashDeepSeek V4.1 FlashGPT-6 Astra
$0.01$0.10

Cost per puzzle → (log scale: each solid line is 10× the last)

1 variant has no positive recorded cost value and cannot appear on a log scale.

View scatter data table

effort ladder

28 of 31 families scored higher at their top effort level than at their lowest. For 7 families, the top level was not their best.

GPT-6 Luna
offminimallowmediumhighxhighmax
GPT-6 Sol
offminimallowmediumhighxhighmax
Claude Opus 5.5
offminimallowmediumhighxhighmax
Grok 4.7
offminimallowmediumhighxhighmax
DeepSeek V4.1 Flash
offminimallowmediumhighxhighmax
GPT-6 Astra
offminimallowmediumhighxhighmax
Qwen3.8 Max
offminimallowmediumhighxhighmax
Muse Spark 1.3
offminimallowmediumhighxhighmax
Gemini 3.8 Flash
offminimallowmediumhighxhighmax
Claude Fable 5.1
offminimallowmediumhighxhighmax
GLM 5.3 Flash
offminimallowmediumhighxhighmax
GLM 5.3
offminimallowmediumhighxhighmax
Qwen3.8 27B
offminimallowmediumhighxhighmax
DeepSeek V4 Pro
offminimallowmediumhighxhighmax
Kimi K3
offminimallowmediumhighxhighmax
GPT-5.6 Sol
offminimallowmediumhighxhighmax
Mistral Medium 3.5
offminimallowmediumhighxhighmax
GPT-5.4
offminimallowmediumhighxhighmax
Gemini 3.1 Pro Preview
offminimallowmediumhighxhighmax
GLM 5
offminimallowmediumhighxhighmax
Kimi K2.5
offminimallowmediumhighxhighmax
GLM 4.7
offminimallowmediumhighxhighmax
Gemini 3 Flash Preview
offminimallowmediumhighxhighmax
GPT-5.2
offminimallowmediumhighxhighmax
DeepSeek V3.2
offminimallowmediumhighxhighmax
Claude Opus 4.5
offminimallowmediumhighxhighmax
gpt-oss-120b
offminimallowmediumhighxhighmax
DeepSeek V3.2 Speciale
offminimallowmediumhighxhighmax
Gemini 3 Pro Preview
offminimallowmediumhighxhighmax
Grok 4.1 Fast
offminimallowmediumhighxhighmax
MiMo V2 Flash
offminimallowmediumhighxhighmax

Reasoning off → on

Claude Sonnet 4.5
offon
View effort data table

detailed model statistics

Accuracy and usage for the selected tier.

Model95% interval5×510×1015×15Hard mode
GPT-6 Astraxhigh
100.0% (89–100%)100%100%100%50%$10.821h 12m
Claude Opus 5.5high
93.3% (79–98%)100%100%80%80%$5.2448m 56s
GPT-5.6 Solhigh
86.7% (70–95%)100%100%60%—$2.2749m 10s
GPT-6 Solhigh
86.7% (70–95%)100%90%70%0%$1.8652m 11s
DeepSeek V4 Prohigh
83.3% (66–93%)100%100%50%0%$2.744h 20m
Gemini 3.8 Flashhigh
83.3% (66–93%)100%100%50%0%$2.4741m 54s
Grok 4.7high
83.3% (66–93%)100%90%60%0%$8.977h 57m
Claude Fable 5.1xhigh
76.7% (59–88%)100%70%60%10%$16.961h 01m
DeepSeek V4.1 Flashhigh
76.7% (59–88%)100%100%30%0%$0.841h 23m
Gemini 3.1 Pro Previewhigh
76.7% (59–88%)100%90%40%—$5.431h 02m
GPT-5.4xhigh
76.7% (59–88%)100%90%40%—$10.202h 38m
Kimi K3max
76.7% (59–88%)100%70%60%0%$14.526h 56m
Muse Spark 1.3xhigh
76.7% (59–88%)100%100%30%0%$3.4049m 53s
GPT-5.2high
66.7% (49–81%)100%80%20%—$10.864h 12m
Qwen3.8 Flashon
66.7% (49–81%)100%100%0%0%$0.736h 16m
Claude Opus 4.5high
60.0% (42–75%)100%50%30%—$23.084h 45m
Qwen3.8 Maxminimal
60.0% (42–75%)100%70%10%0%$4.825h 05m
Gemini 3 Pro Previewhigh
56.7% (39–73%)100%60%10%—$5.201h 20m
Seed 2.1 Turboon
53.3% (36–70%)100%50%10%0%$3.134h 53m
Gemini 3 Flash Previewhigh
50.0% (33–67%)100%40%10%—$1.6858m 45s
GPT-6 Lunamax
50.0% (33–67%)90%60%0%0%$0.391h 50m
DeepSeek V3.2 Specialehigh
46.7% (30–64%)90%30%20%—$0.4413h 02m
MiMo V2.6 Proon
46.7% (30–64%)80%60%0%—$1.006h 23m
Qwen3.8 27Bxhigh
46.7% (30–64%)100%40%0%—$3.617h 12m
DeepSeek V3.2high
43.3% (27–61%)100%20%10%—$0.625h 53m
gpt-oss-120bhigh
43.3% (27–61%)90%40%0%—$0.251h 13m
Kimi K2.5high
43.3% (27–61%)100%30%0%—$2.565h 25m
GLM 5.3 Flashmax
40.0% (25–58%)100%20%0%—$0.668h 29m
GLM 5.3max
36.7% (22–54%)80%20%10%—$4.724h 30m
Grok 4on
33.3% (19–51%)80%20%0%—$15.064h 49m
MiMo V2.6 Flashon
33.3% (19–51%)60%40%0%—$0.3812h 11m
MiniMax M2.5high
33.3% (19–51%)90%10%0%—$1.497h 03m
Qwen3 Next 80B A3B Thinkingon
33.3% (19–51%)100%0%0%—$0.931h 14m
Seed 1.6high
33.3% (19–51%)100%0%0%—$0.832h 08m
Claude Sonnet 4.5on
30.0% (17–48%)90%0%0%—$5.981h 43m
GLM 5high
30.0% (17–48%)70%10%10%—$1.524h 07m
MiniMax M2.1high
30.0% (17–48%)90%0%0%—$0.612h 15m
Mistral Medium 3.5low
30.0% (17–48%)90%0%0%—$7.141h 20m
Kimi K2 Thinkingon
26.7% (14–44%)80%0%0%—$2.984h 31m
Grok 4.1 Faston
23.3% (12–41%)60%10%0%—$0.482h 02m
GLM 4.7on
20.0% (10–37%)60%0%0%—$1.133h 41m
MiMo V2 Flashhigh
20.0% (10–37%)60%0%0%—$03h 02m
OLMo 3.1 32B Thinkon
13.3% (5–30%)40%0%0%—$0.342h 04m
Seed 1.6 Flashhigh
13.3% (5–30%)40%0%0%—$0.201h 14m
Kimi K2off
6.7% (2–21%)20%0%0%—$0.2353m 04s

statistics by grid size

Combined results for the models shown.

5×545 models

88.0%

396/450 solved

1m 55s avg time$0.0265 avg cost
10×1045 models

45.8%

206/450 solved

9m 21s avg time$0.1590 avg cost
15×1545 models

19.3%

87/450 solved

10m 54s avg time$0.2340 avg cost
Hard mode (20×20)14 models

10.1%

14/139 solved

16m 05s avg time$0.7916 avg cost