核心 GUIDE

PATTERNINTERMEDIATE7 分钟阅读

Action Grounding(行动落地)

Action Grounding 要求 Agent 把 Proposed Action 绑定到当前可观察 Environment State、合法 Target 与已验证 Precondition,而不是依赖过期 Assumption 或模型自己的描述。

核心心智模型

Grounding 会闭合 Perception → Action Loop。执行高后果 Action 前,Agent 应知道正在作用于哪个 Object/State、该状态如何被观察、Precondition 是否仍成立。

为什么重要

Computer-use / Tool-using Agent 即使 Plan 正确,也可能点错 Element、操作错 Account 或依据 Stale Page 执行。Natural-language Intent 不足以支撑可靠执行;Runtime 需要 Identifier、Observation 与 Post-action Verification 把模型 Decision 连接到真实环境。

01

Observe、Bind、Act、Verify

在接近 Action 的时间点获取 Fresh State,把 Intended Target 绑定到 Stable Identifier 或 Validated Selector,检查必要 Precondition,再通过 Scoped Capability 执行。随后重新 Observe Environment,把 Expected State Transition 与实际结果比较;Observation 已变化时应该 Replan,而不是强行执行旧 Action。

02

示例:Agent 点击了错误的 Delete Button

Dashboard 多行都有相同 Delete Label。模型 Reasoning 里描述的是正确 Customer,但 Table 在执行前重新排序,Position Selector 指向另一行。Grounding 会绑定 Customer Stable Row ID,执行前再次核对 Displayed Account Name,并在执行后确认只有目标 Record 消失。

常见失败模式

  • 基于可能已经 Stale 的 Screenshot/State Observation 执行。
  • 有 Stable Identifier 时仍依赖 Visual Position 或 Ambiguous Label。
  • Tool Response 显示 Success 就假设 Intended Environment State 已改变。

工程启发

  • 高后果 Action 前刷新关键 State。
  • 把 Action 绑定到 Stable Identifier 与显式 Precondition。
  • 验证 Post-action World,而不只是 Tool Call Status。

关键结论

  1. 01正确 Intent 不保证 Action Grounding 正确。
  2. 02Grounding 把 Model Decision 连接到 Observable Environment State。
  3. 03Post-action Verification 属于 Grounding Loop。

阅读记录

未打开Practice 尚未完成

这里只记录你真实做过的动作,不代表掌握、熟练或认证。

用于这些学习路径

这个 Concept 会在多个 canonical 学习路径中复用。

来自 Knowledge Graph 的相关知识点

这些关系直接来自 canonical graph,不维护第二套 Guide 分类体系。

行动前观察环境PREREQUISITE