【AI 核心深度 M4-060】解释连续批处理(continuous batching)与静态批处理的差异。(Continuous Batching vs. Static Batching in LLM Serving)深度数理推导与工程落地解析

所属模块:M4 · 序列与 Transformer (Sequences & Transformers) | 专题分类:KV Cache 与推理优化 (KV Cache & Inference Optimizations) | 难度等级:Medium

一、核心一句话结论 (One-Sentence Summary)

静态批需等整批完成才能换新请求(GPU 空转);连续批在每步迭代后动态插入/移除请求,大幅提升吞吐。

ADVERTISEMENT · 赞助推荐

Continuous batching operates at the iteration level to dynamically insert arriving requests and evict completed requests at every decoding step, eliminating the massive bubble overhead of static sequence-level batching.

二、核心考点要义 (Key Insights)

  • 📌 静态批:批内请求长度不一 → 短请求完成后 GPU 空转
  • 📌 连续批:每个 decode step 后重新组批,空闲槽立即补新请求
  • 📌 吞吐可提升数倍,是现代推理引擎的标配

English Insights:
– Static batching: requests are grouped at the sequence level and must wait for the longest sequence to finish, causing massive GPU idle bubbles and high tail latency
– Continuous (iteration-level) batching: the batch is reconstituted at every single token step; finished sequences exit immediately and new prompts enter without waiting
– Enables $3times$ to $10times$ higher serving throughput and drastically reduces average Time-to-First-Token (TTFT)

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$text{static}: text{wait for longest};qquad text{continuous}: text{refill slots every iteration}$$

数学机理:静态批处理(static batching) 把一批请求一起提交、一起等待完成:因为不同请求的输出长度不同,整个批的耗时由最长的那个决定;短请求完成后其占用的资源空闲但无法被利用(必须等整批结束)。故 GPU 利用率 = (总有效计算)/(批耗时 × 资源),在长度差异大时利用率很低(可能 <50%)。连续批处理(continuous batching,又称 iteration-level scheduling,Yu 等 2022 Orca) 的核心:在每个 decode step 之后重新组批——已完成(生成 EOS 或达到上限)的请求立即移出、腾出的槽位立即补入等待队列中的新请求。这样 GPU 几乎始终满负荷,吞吐大幅提升(Orca 报告 2~36 倍,取决于长度分布)。关键前提——(a) KV cache 需能动态分配/释放(PagedAttention 提供块级管理);(b) 注意力需支持批内不同长度(变长序列的批处理,用 mask 与 ragged tensor);(c) 调度器需在每步决定’谁进谁出’(如按 FCFS、或按优先级/公平性)。与 chunked prefill 的结合——新请求的 prefill(长 prompt)会占用大量算力、阻塞 decode;故现代引擎把 prefill 切块(chunked)并与 decode 混批,进一步优化延迟与吞吐。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Mathematical Mechanism: Static Batching Bubble Problem: Consider a batch of $B$ requests with prompt lengths $L_{i}^{text{prompt}}$ and generation lengths $L_{i}^{text{gen}}$. The static batch execution time is determined by $max_i (L_i^{text{prompt}} + L_i^{text{gen}})$. Requests that finish early must output padding tokens until the slowest request finishes. The GPU execution efficiency $eta$ is: $$eta = frac{sum_{i=1}^B (L_i^{text{prompt}} + L_i^{text{gen}})}{B times max_i (L_i^{text{prompt}} + L_i^{text{gen}})}$$ In conversational workloads with high length variance (e.g., lengths ranging from 20 to 2048 tokens), $eta$ can drop below 20%, wasting >80% of GPU compute and memory bandwidth on padding. Continuous Batching Mechanism: Execution is scheduled at the granularity of a single decoding iteration $t$. At step $t$, the batch tensor $X_t in mathbb{R}^{B_t times 1 times d}$ contains active tokens from $B_t$ ongoing sequences. When sequence $k$ outputs an EOS token, its state is immediately evicted, releasing its KV cache. A pending request from the queue can immediately execute its prefill phase (or first decode step) in step $t+1$. The effective utilization approaches $eta approx 1.0$, bounded only by KV cache VRAM limits and scheduling overhead.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

深度剖析与工程权衡:① 为什么这是’吞吐’而非’延迟’优化——连续批提升的是整体吞吐(单位时间处理的请求数);单请求延迟可能略增(因为要与其他请求竞争)。故它适合’高负载服务’。② 调度策略的权衡——(a) FCFS(先来先服务)公平但可能被长请求阻塞;(b) 优先级调度(如区分交互式与批处理);(c) 公平调度(如 vLLM 的’公平共享’,防止某请求长期占用);调度器是推理引擎的核心竞争力之一。③ 与 PagedAttention 的协同——连续批需要’动态分配与回收 KV 块’,而 PagedAttention 的块级管理正好提供这一能力;两者是配套技术(vLLM 同时实现)。④ 变长注意力的实现——批内序列长度不同,注意力 kernel 需支持’每个序列有自己的 KV 长度’(用块表与变长 mask);这是工程复杂度所在。⑤ 与投机解码的交互——投机解码使每个请求每步推进的 token 数不同(取决于接受长度),进一步增加调度的动态性。⑥ 面试要点——被问’如何提升推理吞吐’,应给出’连续批处理(动态组批)+ PagedAttention(动态 KV 管理)+ chunked prefill(混批)‘的组合,并说明’静态批的浪费来自最长请求’;能提到调度策略(FCFS/优先级/公平)是深度理解的标志。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

Deep Dive & Engineering Trade-offs: ① PagedAttention Dependency: Continuous batching causes dynamic, non-contiguous memory allocations as sequences grow and terminate at different rates. Without virtual memory paging (PagedAttention), dynamic memory allocation suffers severe external fragmentation. ② Preemption and Eviction: When VRAM is exhausted due to growing KV caches in the active batch, the engine must either preempt (swap to CPU memory) or recompute (drop KV cache and re-run prefill later) low-priority requests. ③ Mixed Batch Scheduling: Co-scheduling prefill tokens (GEMM) and decode tokens (GEMV) in the same iteration requires chunked prefill to prevent prefills from starving ongoing decode latencies. ④ System Architecture: Pioneered by Orca and popularized by vLLM, TensorRT-LLM, and TGI, continuous batching is the universal de facto standard in production LLM inference engines. ⑤ Interview Strategy: Contrast sequence-level scheduling with step-level scheduling, draw the execution timeline diagram showing how bubbles are eliminated, and explain how it ties into KV cache management.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 以为静态批与连续批只是’批大小不同’
  • ⚠️ 忽略连续批对 KV 动态管理能力的依赖

English Pitfalls:
– Believing static batching padding does not consume memory or compute (padding tokens still undergo matrix multiplications)
– Overlooking that continuous batching requires dynamic memory management like PagedAttention to prevent VRAM fragmentation
– Failing to explain what happens when memory runs out during continuous generation (preemption/recomputation)

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 为什么连续批能大幅提升吞吐?
  2. How does an engine handle VRAM exhaustion when all active requests in a continuous batch expand their KV caches?
  3. 连续批与 PagedAttention 的关系?
  4. What is the difference between iteration-level scheduling in Orca and vLLM?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:KV Cache 显存占用公式、Prefill/Decode 阶段与 PagedAttention (KV Cache Memory, Prefill/Decode & PagedAttention)
  • 🗺️ 知识图谱模块:AI 基础设施工程导图

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M4-060) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.