跳到正文
原文
arXiv cs.AI· Wensen Wu·· 11 小时前精选AI 评分78

Kepler 提出可审计世界模型并在 ARC-AGI-3 公开集取得满分

Kepler: Auditable World Models for ARC-AGI-3

AI 导读

Kepler 在 ARC-AGI-3 的 25 个公开游戏取得 server-verified 100.00 RHAE,且未做 per-game model selection 或 score-conditioned reruns。它是一个开源 harness,把假设表示为可执行世界模型,并用 retrospective transition checks 和 conditional prediction checks 验证;关键不是满分本身,而是公开集分数在交互 agent 评测中的判别力问题。

证据细节有两组。第一组是效率:183 个完成关卡中 181 个的最终 Opus attempt 使用的动作数不超过对应 median-human baseline,保留的 board runs 共 8,256 个 environment actions,其中 7,292 个发生在 scored levels。第二组是成本与验证:保留的 local provider-session records 显示 858.0 million tokens、97.37% cache reads,按 2026 年 9 月 1 日 API list-equivalent rates 成本为 $777.72;论文同时报告 source-code leakage 造成无效完美运行、控制条件下 agents 重建 removed harness、autonomous repair 掩盖 broken planner 三类失败。

论文被 NeurIPS 2026 的 Interpreting Agent Behavior workshop 接收,代码和公开 traces 可用。事实是,在最终 Claude Opus 5 和 GPT-5.6 Sol boards 中,50 个 game-model cells 有 48 个达到 100,且单游戏观察案例显示 animation frames 包含 settled text grids 中没有的任务相关信息。推断是,对 ARC-AGI-3 这类需要从观察推断规则和目标的任务,公开集分数若脱离 first-attempt、cost-conditioned 和 verification-aware 指标,容易把模型能力、harness 漏洞和评测泄漏混在一起。猜测是,后续 agent 基准会更强调可审计轨迹、成本约束和首次尝试,而不是只报告公开集最高分。

推荐理由

它把公开集满分拆到动作数、模型成本、缓存读取和验证失败,能校正对智能体真实能力与评测可信度的判断。

来源:arXiv cs.AI · arxiv.org