AI Resonance · 阅读最新日报 · AI 入门推荐

2026-09-25 · America/Los_Angeles · 论文 · #14

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

Scientific discovery begins where known problems end. There, AI systems must engage in exploration: framing hypotheses, designing experiments, and iterating on the results. However, evaluating this ability is difficult: (1) how to verify whether a genuinely new hypothesis holds, and (2) how to determine whether a system has discovered it through exploration or merely recalled related knowledge from pre-training data. To this end, we introduce ExplorationBench, which turns the wicked problem of evaluating scientific exploration into a concrete and tractable framework built on verifiable Alien…

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

热度 35.2 / 100;排名与评分保留该期记录。