2026-09-18 · America/Los_Angeles · 社区动态 · #15
Open-weight models match Fable 5 on hard LiveCodeBench at a tenth of the cost
Training-free manager-worker scaffold: fresh instances of one model decompose the problem and coordinate. On hard LiveCodeBench, orchestrated models approach or match Fable 5 at a fraction of the cost. FlashNext runs $5.76 a pass against Fable 5's $61.11. - Qwen3.8-FlashNext: 84.2% --> 93.0% - Qwen3.8-27B: 69.2% --> 92.4% - GPT-5.6-Terra: 80.8% --> 88.0% - GPT-5.6-Luna: 70.4% --> 81.2% - Claude Fable 5: 90.4% single call (no harness) https://github.com/slee-persis/GVS5H https://arxiv.org/pdf/2608.26480
热度 45.3 / 100;排名与评分保留该期记录。