AI Resonance · 阅读最新日报 · AI 入门推荐

2026-10-02 · America/Los_Angeles · 社区动态 · #10

Evaluating 14 LLMs as a visual coding agent: re-rendering, a deterministic judge, and the provider quirks that skewed my first results

I built a bench for a photo-to-Blender agent (the model writes and runs Blender Python to rebuild a photo as an editable scene) and ran 14 models through it. The design choices that mattered most: Don't score what the model hands in. The bench opens each delivered scene file and renders it itself, at the photo's aspect ratio, max 1,100 px, 96 samples. No LLM judge. Estimated silhouette overlap (30%), edge F1 (40%), multiscale RGB error (15%) and mesh health (15%: non-manifold and open edges, duplicate faces, inconsistent winding, degenerate triangles), mapped through calibrated…

Evaluating 14 LLMs as a visual coding agent: re-rendering, a deterministic judge, and the provider quirks that skewed my first results

热度 45.3 / 100;排名与评分保留该期记录。