所属模块:
M6 · 多模态与生成模型 (Multimodal & Generative Models)| 专题分类:引导与采样 (Guidance & Fast Sampling (CFG / DDIM))| 难度等级:Hard
一、核心一句话结论 (One-Sentence Summary)
CFG 需条件与无条件两次前向(成本 ×2);可用 CFG distillation、批量并行或近似合并来降低成本。
Classifier-Free Guidance doubles inference compute by evaluating conditional and unconditional branches at every step, motivating acceleration techniques like CFG distillation, dynamic guidance scheduling, and negative-prompt caching.
二、核心考点要义 (Key Insights)
- 📌 CFG 每步需两次前向(有条件 + 无条件)
- 📌 成本 ×2 是扩散推理的主要开销之一
- 📌 优化:CFG distillation(合并为一次)、批量并行、部分步 CFG、更好求解器
English Insights:
– The 2x compute burden: CFG doubles forward-pass FLOPs ($,mathcal{O}(2 times N_{text{steps}} times text{FLOPs}_{text{model}}),)$, making guidance the primary latency bottleneck in diffusion serving
– Dynamic CFG scheduling: applies guidance exclusively during early structural timesteps ($t in [0.5T, T]$) and reverts to single-pass conditional generation for late detail steps, saving 35-50% compute
– CFG distillation (Auto-Guidance): distills guided two-stream generation into a single-forward-pass student network, recovering $2times$ serving throughput with zero quality loss
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$hatepsilon=epsilon_theta(x_t,varnothing)+s(epsilon_theta(x_t,c)-epsilon_theta(x_t,varnothing));qquad text{cost}times2$$
数学机理:CFG 的公式与成本——ε̂=εθ(x_t,∅)+s·(εθ(x_t,c)−εθ(x_t,∅)),其中 εθ(x_t,c) 是条件预测、εθ(x_t,∅) 是无条件(空条件)预测。故每步需两次前向(条件 + 无条件),推理成本约为无 CFG 的 2 倍(在’采样步数’相同时)。这是扩散推理的主要成本之一(因为采样本身已需数十步)。优化手段——(1) CFG Distillation——训练一个学生模型,其单次前向的输出直接等于 CFG 的组合结果(ε̂):(a) 学生把 s 作为输入(可调引导强度),用教师(带 CFG)的输出监督;(b) 或用’CFG++’类方法(改进引导的数值形式使其更易蒸馏)。收益——成本减半(2×→1×)。(2) 批量并行——把’条件’与’无条件’两次前向合成一个 batch(batch 维拼接)一次前向完成;收益——不减少计算量(FLOPs 相同),但减少 kernel 启动与内存往返,故墙钟时间可能显著降低(尤其小 batch 时);这是最简单的优化(一行代码)。(3) 近似/部分方法——(a) 只对部分步用 CFG(如前几步用、后几步不用);(b) 用历史步推断无条件预测;(c) CFG++ 的改进形式(减少前向次数)。(4) 减少步数——用更好的求解器(DPM-Solver 10~20 步),从根本上降低’每步成本 × 步数’。(5) 潜空间 + 量化——降低单次前向的成本(见 LDM 与量化题)。与’条件注入方式’的关系——CFG 需要’无条件’前向,故条件注入方式需支持’空条件’(如把文本嵌入设为空/零、把类别设为 null);这是实现细节。与’引导强度 s’的关系——s=1 时公式退化为 εθ(x_t,c)(只需一次前向);故’不需要引导’的场景可省一半成本。实践建议——(a) 必做:批量并行(零成本);(b) 值得做:CFG distillation(成本减半,需训练);(c) 可考虑:部分步 CFG、更好的求解器。度量——(a) 每张图的延迟;(b) 质量(FID/CLIP-score);(c) 蒸馏后 s 的可调性是否保持。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Mathematical Mechanism: 1. Latency & FLOP Quantification: Standard unguided diffusion requires $N$ forward passes: $$text{Compute}_{text{no-CFG}} = N cdot mathcal{C}_{text{model}}$$ Under Classifier-Free Guidance with conditional $c$ and unconditional $emptyset$: $$text{Compute}_{text{CFG}} = N cdot big( mathcal{C}_{text{model}}(x_t, c) + mathcal{C}_{text{model}}(x_t, emptyset) big) = 2 N cdot mathcal{C}_{text{model}}$$ Doubling memory bandwidth requirements and cutting serving requests-per-second (RPS) in half. 2. Dynamic Timestep Guidance Gating: Empirical studies show that CFG establishes global semantic layout and composition during high-noise steps ($t > t_{text{cutoff}}$), while low-noise steps ($t le t_{text{cutoff}}$) simply refine local textures where guidance is redundant: $$s(t) = begin{cases} s_{text{target}} & text{if } t > t_{text{cutoff}} quad (text{dual pass, } 2times text{ cost}) \ 1.0 & text{if } t le t_{text{cutoff}} quad (text{single pass, } 1times text{ cost}) end{cases}$$ Setting $t_{text{cutoff}} = 0.4 T$ eliminates $40%$ of unconditional forward evaluations. 3. Guidance Distillation Objective (Meng et al., 2023): Trains a single student model $hat{epsilon}_theta(x_t, c, s)$ to directly output the guided trajectory: $$mathcal{L}_{text{distill}} = mathbb{E}left[ big| hat{epsilon}_theta(x_t, c, s) – big( (1-s)epsilon_{theta^*}(x_t, emptyset) + s epsilon_{theta^*}(x_t, c) big) big|^2 right]$$ Permitting 1 forward pass per step at arbitrary guidance scale $s$.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
深度剖析与工程权衡:① ‘CFG 的成本 ×2’是扩散推理的常识——面试中能指出这一点(而非只说’CFG 提升质量’)是深度理解的标志。② ‘批量并行是零成本优化’——它不减少 FLOPs 但降低墙钟时间(减少 kernel 启动与内存往返);应默认使用。③ ‘CFG distillation 的取舍’——成本减半但需训练,且需保持 s 的可调性。④ ‘s=1 时只需一次前向’——因为公式退化为条件预测;故无需引导时可省一半。⑤ ‘与少步采样的叠加’——减少步数(10~20 步)与 CFG distillation(减半)叠加,可把总成本降到 1/10 以下。⑥ 面试要点——被问’CFG 的成本’,应给出’每步两次前向 → 成本 ×2‘与’优化(批量并行零成本 / CFG distillation 减半 / 部分步 CFG / 更好求解器)‘;能指出’批量并行不减少 FLOPs 但降低墙钟时间’是深度理解的标志。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Deep Dive & Engineering Trade-offs: ① Batched Forward Pass vs Sequential Passes: In GPU serving engines, conditional and unconditional passes are concatenated into a single batch of size $2B$: $[x_t, x_t]$ paired with $[c, emptyset]$. While this maximizes tensor core utilization, it doubles activation memory, which can force serving systems to reduce max concurrency or trigger VRAM out-of-memory errors on large models (DiT-3B). ② Negative-Prompt Embedding Caching: Because the unconditional/negative prompt embedding $c_{text{neg}}$ is identical across all requests or fixed across sampling steps, its text encoder projection is pre-computed and stored globally in GPU memory, avoiding redundant text encoder calls. ③ PAG (Perturbed Attention Guidance): A recent alternative that replaces the unconditional pass with a structurally perturbed attention pass (zeroing out attention self-weights), enhancing generation quality without text prompts. ④ One-Step Distillation Paradigms: In ultra-fast architectures like SDXL-Turbo and Flux-Schnell, models are distilled using adversarial objectives directly into single-pass generators operating without CFG ($s=1.0$), completely bypassing the 2x tax. ⑤ Interview Strategy: Quantify the $2times$ latency overhead, explain dynamic CFG scheduling over time $t$, derive the guidance distillation objective, and discuss batched GPU execution trade-offs.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 忽略 CFG 的双倍前向成本
- ⚠️ 不做批量并行(白白浪费墙钟时间)
English Pitfalls:
– Applying full CFG across all 50 timesteps when late high-frequency steps gain zero perceptual benefit from unconditional subtraction
– Executing conditional and unconditional forward passes sequentially in Python loops rather than concatenating them into a single $2B$ batch
– Assuming guidance distillation requires re-generating training datasets from scratch; distillation uses the pre-trained teacher on the fly
六、高频深度面试追问与预测 (Follow-Up Questions)
- CFG distillation 怎么做?
- Why is Classifier-Free Guidance predominantly impactful during high-noise early timesteps rather than late low-noise refinement steps?
- 批量并行能省时间吗?
- How does Guidance Distillation compress the two-stream CFG computation into a single forward pass per sampling step?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
扩散引导与加速采样:Classifier-Free Guidance (CFG) 与 DDIM 确定性采样(Classifier-Free Guidance (CFG) & Accelerated DDIM Sampling) - 🗺️ 知识图谱模块:
多模态与扩散模型导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。