2026-09-29 · America/Los_Angeles · 论文 · #8
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and adherence to task requirements. This leaves the policy without a learning signal that favors clean, targeted implementations over those containing unnecessary or out-of-scope changes. We introduce GAGAR, a framework for quality-aware credit redistribution in code agent RL. Built on dynamic…
热度 53.5 / 100;排名与评分保留该期记录。