AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-24 · America/Los_Angeles · 论文 · #13

WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents

AI research agents need reliable knowledge of how their experiments change outcomes. We introduce WhatWorkedBench to measure experimental understanding, the accuracy of predictions about component changes after budgeted experimentation. Agents inspect code, select measurements, and submit a response surface, a table predicting scores for every configuration of component settings. Exhaustive CPU execution supplies reference effects for changing each component while holding the others fixed. These effects capture combinations of changes across 36 tasks from 30 data sources and 8 workflow…

WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents

热度 38.0 / 100;排名与评分保留该期记录。