2026-10-01 · America/Los_Angeles · 论文 · #15
More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models
Direct-decision models turn text into low-latency structured labels and scores, making them attractive for classification and automatic evaluation. Yet reliability requires more than accuracy: a model must also use the ordinal decision scale supplied by the user faithfully. We analyze JEV~1.13 and three open KEV models. Our investigation begins with ANLI, where JEV assigns 38.8\% of all predictions and 51.3\% of errors to Neutral despite 74.95\% accuracy, nearly balanced gold labels, and balanced candidate positions. Across 36 ordinal datasets, final decisions use only 67--76\% of the…
热度 51.3 / 100;排名与评分保留该期记录。