2026-10-02 · America/Los_Angeles · 社区动态 · #18
Adding memory to search instead of sampling in reward maximization tasks [R]
I am one of the authors of FLEET - an algorithm that enhances Best-of-N generation by attributing external rewards to particular tokens and then uses MCTS to adjust logits during the next run. I find it rather funny that most of the tasks where repetitive sampling is widely used are based on reward maximization, yet it is not aware of that reward. Tuning the sampling parameters allows to make the process more efficient, but it is still a blind search. We propose a way to make generation aware of previous rewards with solutions on how to attribute reward to the completion and how to use this…
热度 44.3 / 100;排名与评分保留该期记录。