2026-10-04 · America/Los_Angeles · 社区动态 · #16
Nonobench: an open benchmark of 49 LLMs on nonogram puzzles, public and open source [P]
Nonobench measures how well LLMs solve nonograms (picross). Each model gets the row and column clues once and returns the full grid. No tools, one attempt per puzzle. Method: - Standard mode: 30 puzzles from 5x5 to 15x15 (from the Nonograms dataset by Moyà-Alcover, CC BY 4.0). - Hard mode: ten random 20x20s, each checked to have a single solution. Five can't be solved by line logic alone. Random fills avoid picture puzzles that models can guess. - 130 variants across reasoning effort levels, run through OpenRouter and pinned to each lab's own endpoint where possible. Results: - Solve rates…
热度 37.1 / 100;排名与评分保留该期记录。