【AI 核心深度 M5-091】解释 Agent 的规划(planning)与任务分解。(Agent Planning and Hierarchical Task Decomposition)深度数理推导与工程落地解析

所属模块:M5 · NLP 与大语言模型 (NLP & Large Language Models) | 专题分类:Agent 与工具调用 (Agents & Tool Use) | 难度等级:Medium

一、核心一句话结论 (One-Sentence Summary)

先把目标分解为子任务(显式计划),再逐步执行;比’边想边做’更能处理长任务与依赖关系。

ADVERTISEMENT · 赞助推荐

Planning equips agents with the ability to decompose complex objectives into structured, dependency-aware subtasks prior to execution, preventing the myopic drift and combinatorial trial-and-error common in purely greedy ReAct systems.

二、核心考点要义 (Key Insights)

  • 📌 显式计划:先分解为子任务(含依赖关系)
  • 📌 对比 ReAct 的’边想边做’:计划更全局、少迷路
  • 📌 需支持重规划(执行中发现问题则调整计划)

English Insights:
– Decomposition taxonomy: Plan-and-Execute (global upfront planning), Tree-of-Thought / LATS (tree search over action spaces), and Hierarchical Planning (nested sub-agents)
– Granularity trade-off: overly coarse plans provide zero execution guidance, while overly granular plans break immediately upon dynamic environmental surprises
– Dynamic replanning: continuous state monitoring that triggers targeted plan revision when subtask execution fails or uncovers contradictory facts

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$text{plan}: gto{s_1,dots,s_k}totext{execute};qquad text{replan on failure}$$

数学机理:规划的必要性——纯 ReAct 是’贪心的’(每步只看眼前),在长任务中容易 (a) 迷路(偏离目标);(b) 重复(忘记做过什么);(c) 无法处理依赖(子任务有先后顺序)。规划(planning) 的做法:先让 LLM 把目标 g 分解为子任务 {s_1, …, s_k}(显式的计划),再逐步执行;执行中可重规划(replan)。分解的粒度——(a) 太粗(如’完成任务’)——无指导作用;(b) 太细(如’按第一个键’)——过于机械、无法适应意外;合适的粒度是’有意义的、可独立验证的步骤’(如’搜索 X 的定义’、’比较 A 与 B’)。依赖关系——子任务之间可能有顺序依赖(必须先 A 后 B)或可并行;显式记录依赖可 (a) 并行化可并行的部分、(b) 避免顺序错误。架构模式:(1) Plan-and-Execute——先生成完整计划,再逐步执行(执行器可以是简单 Agent);优点:全局视野、成本可控(计划一次);缺点:计划可能不适用(环境变化时需重规划)。(2) ReAct(无显式计划)——边想边做;优点:灵活适应;缺点:长任务易迷路。(3) 混合——先粗粒度计划,执行时用 ReAct 处理每个子任务(’计划 + 反应’);这是当前的推荐做法。(4) 层级规划——递归分解(大计划 → 子计划 → 原子步骤)。(5) 搜索式规划——在动作空间搜索(如 LATS 的 MCTS);更强但更贵。重规划(replan)——执行中若发现 (a) 子任务失败、(b) 环境变化、(c) 新信息与假设不符,则重新生成计划(可保留已完成的步骤)。与’工作记忆’的关系——计划本身就是工作记忆的核心内容(显式记录’要做什么、做到哪了’)。评估——(a) 计划的可行性(子任务是否可执行);(b) 计划的完整性(是否覆盖所有必要步骤);(c) 执行成功率。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Mathematical Mechanism: 1. Plan Formalism: Given root objective $G$, a planner generates an acyclic directed dependency graph $mathcal{G} = (mathcal{V}, mathcal{E})$, where vertices $v_i in mathcal{V}$ represent subtasks and directed edges $(v_i, v_j) in mathcal{E}$ enforce sequential execution constraints ($v_i prec v_j$): $$P(mathcal{G} mid G) = prod_{i=1}^k P(v_i mid G, text{Parents}(v_i))$$ Independent vertices $(text{InDegree}(v) = 0)$ can be dispatched for parallel execution. 2. Language Agent Tree Search (LATS): Integrates Monte Carlo Tree Search (MCTS) into the planning loop: $$text{UCT}(s, a) = Q(s, a) + c sqrt{frac{ln N(s)}{N(s, a)}}$$ where value function $Q(s, a)$ is scored via model self-reflection or programmatic test execution, balancing exploratory action trajectories against exploitation of validated paths.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

深度剖析与工程权衡:① ‘显式计划’解决长任务的’迷路’问题——它是 Agent 从’玩具’到’可用’的关键;面试中能对比’Plan-and-Execute vs ReAct’的适用场景是加分。② ‘计划太细’是常见错误——过度细化的计划无法适应执行中的意外(因为每步都被固定);故计划应留出’执行时的灵活空间’(粗粒度计划 + 细粒度执行)。③ ‘重规划’是必需能力——现实任务中计划必然需要调整(工具失败、信息不符);故架构需支持’边执行边修正计划’。④ ‘成本结构’的差异——Plan-and-Execute 的计划只需一次 LLM 调用(便宜),执行阶段可用简单模型;ReAct 每步都需推理(更贵)。故对’结构化的长任务’,Plan-and-Execute 更经济。⑤ 与’多智能体’的关系——’规划者 + 执行者’的多 Agent 架构是规划思想的一种实现;但也可在单 Agent 内完成(先生成计划再执行),无需真正的多 Agent(避免通信开销)。⑥ 面试要点——被问’Agent 如何规划’,应给出’显式分解(粒度适中)+ 依赖关系 + 重规划 + Plan-and-Execute 与 ReAct 的对比与混合‘,并强调’计划太细反而有害‘与’长任务需要显式计划防迷路‘;这是 Agent 架构类问题的深度回答。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

Deep Dive & Engineering Trade-offs: ① Plan-and-Solve vs ReAct: Pure ReAct is greedy and myopic ($T$ local decisions); it excels in dynamic, exploratory environments but frequently loses global coherence on multi-step workflows. Plan-and-Execute generates a global roadmap first, reducing token costs (planner runs once, cheap executors handle subtasks), but is fragile to unexpected environmental state changes. ② Hybrid Architecture (Plan-then-ReAct): The industry standard pattern generates a coarse-grained high-level plan (3-6 key milestones), where each milestone is executed by an autonomous ReAct loop. If a milestone fails, control returns to the planner for dynamic replanning. ③ Granularity Calibration: The cardinal rule of agent planning is: ‘Plan milestones, not keystrokes’. Over-specifying atomic steps (e.g., ‘Step 1: click button X; Step 2: enter text Y’) causes the entire plan to fail if the UI slightly differs. Plans should define verifiable outcomes (e.g., ‘Subtask: retrieve and parse user billing history’). ④ Explicit Execution Scratchpads: Maintaining an updated checklist (e.g., `[x] Subtask 1, [/] In Progress Subtask 2, [ ] Subtask 3`) in the active context dramatically suppresses task forgetting and redundant tool calls. ⑤ Interview Strategy: Contrast upfront planning with reactive stepping, draw the Plan-then-ReAct hybrid architecture, formulate the replanning trigger mechanism, and discuss MCTS-based search (LATS) for complex coding agents.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 计划过于细化(无法适应意外)
  • ⚠️ 不支持重规划(环境变化时计划失效)

English Pitfalls:
– Generating hyper-granular, rigid upfront plans that break upon the very first minor API or environment mismatch
– Running pure upfront planning without implementing dynamic replanning triggers when an intermediate subtask fails
– Using pure greedy ReAct on 20+ step workflows without an explicit goal tracking scratchpad, causing severe task drift

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 计划太细或太粗有什么问题?
  2. Why does a hybrid Plan-then-ReAct architecture consistently outperform both pure ReAct and pure Plan-and-Execute?
  3. Plan-and-Execute 与 ReAct 如何结合?
  4. How does Language Agent Tree Search (LATS) adapt Monte Carlo Tree Search to discrete reasoning and action spaces?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:智能体系统架构:ReAct 循环、Function Calling、反思记忆与状态机控制 (AI Agents: ReAct Paradigm, Function Calling & Finite State Machines)
  • 🗺️ 知识图谱模块:AI 应用与 Agent 拓扑导图

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M5-091) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.