2026-10-02 · America/Los_Angeles · 论文 · #16
Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning progress. We study this behavior through advantage energy, the sum of a reward's squared advantages over a batch. Under idealized GDPO normalization, we show that this energy is proportional to active-group density: the fraction of rollout groups in which the reward provides nonzero relative…
热度 49.8 / 100;排名与评分保留该期记录。