AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-23 · America/Los_Angeles · 论文 · #2

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

LLM agents increasingly work on long-horizon tasks, and the decisions they make along the way, such as which hypothesis to test or which implementation to build on, determine the outcome of the whole run. Making these decisions well is becoming a key capability for both engineering and research agents. We refer to the ability to make good long-horizon decisions as the taste of an agent. While existing benchmarks measure the end-to-end success of agents on long-horizon tasks, none of them measures the taste of an agent. To address this problem, we build Taste-Bench, a benchmark of taste…

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

热度 57.2 / 100;排名与评分保留该期记录。