AI Resonance · 阅读最新日报 · AI 入门推荐

2026-10-01 · America/Los_Angeles · 论文 · #9

False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward improves without a matching gain in external correctness. A post-hoc audit against source evidence shows co-cheating growing more severe over successive rounds of self-evolution, with pseudo-label correctness stagnating or declining even as the in-loop training signal improves. The most direct…

False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

热度 55.5 / 100;排名与评分保留该期记录。