AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-24 · America/Los_Angeles · 论文 · #5

RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling

Reinforcement learning (RL) is vital for optimizing video generation models, with a robust reward model (RM) serving as the cornerstone. However, existing video reward models often produce unstable scalar scores because they directly map complex, subjective video quality into a single score without explicit evaluation criteria. This leads to scalar drift, where the scoring scale collapses or shifts across different prompts, making the reward unreliable for RL. Drawing inspiration from professional human annotation engineering, we address this problem with RewardVerse, a rubric-based video…

RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling

热度 45.9 / 100;排名与评分保留该期记录。