2026-10-01 · America/Los_Angeles · 社区动态 · #15
LLMs that push back on a wrong user still accept the same wrong answer from a "verified source" - NeurIPS 2026 [R]
I'm one of the authors. We kept seeing models that hold their ground when the user insists on a wrong answer, yet change their answer when the same claim is framed as coming from a "verified source". We wanted to measure how often this happens and check whether the model represents the two cases differently. We call the effect Authority Bias. Why we think it matters. Standard sycophancy evals apply pressure through the user, so a model can pass them while still being easy to mislead through search results, retrieved documents and tool outputs. Another reason is with current AI research…
热度 46.1 / 100;排名与评分保留该期记录。