所属模块:
M6 · 多模态与生成模型 (Multimodal & Generative Models)| 专题分类:视觉编码器 (Vision Encoders (ViT / ConvNeXt))| 难度等级:Easy
一、核心一句话结论 (One-Sentence Summary)
CNN 有局部性/平移等变先验(样本效率高);ViT 偏置弱(需大数据),但数据充足时上限更高。
CNNs embed strong inductive biases of locality and translation equivariance for high sample efficiency on small datasets, whereas ViTs discard spatial priors in favor of global attention, achieving superior performance ceilings when trained on massive data.
二、核心考点要义 (Key Insights)
- 📌 CNN:局部性 + 平移等变 + 权重共享(强先验)
- 📌 ViT:全局注意力,无空间先验(需位置编码)
- 📌 交叉点:ImageNet-1k CNN 优;JFT-300M ViT 优
English Insights:
– CNN inductive biases: strictly hardcodes locality (small receptive field) and translation equivariance (shared sliding kernel weights)
– ViT inductive openness: global self-attention computes pairwise token affinities with zero hardcoded spatial priors, requiring explicit positional encodings
– Data regime crossover: on small-scale datasets (ImageNet-1k), CNNs outperform ViTs; on web-scale data (JFT-300M, LAION-5B), ViTs achieve higher capacity and superior generalization
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$text{CNN}: text{locality}+text{equivariance};qquad text{ViT}: text{weak prior}Rightarrow text{data-hungry}$$
数学机理:归纳偏置(inductive bias) 的差异(详见 M4 的’注意力归纳偏置’题)。(1) CNN——内置 (a) 局部性(每个输出只看 k×k 邻域)、(b) 平移等变性(同一卷积核滑动使用)、(c) 权重共享;这些先验在图像上非常正确(局部纹理、平移不变),故 CNN 在小数据上样本效率高、训练稳定。(2) ViT——用全局自注意力(每个 patch 与所有 patch 交互)+ 位置编码;无局部性与平移等变先验(注意力对 patch 的排列是等变的,故需位置编码注入顺序)。后果——ViT 需要更多数据才能学到这些结构(否则在小数据上欠拟合);但一旦数据充足,ViT 的灵活性(可学习任意长程依赖)使其上限更高。实证(ViT 论文)——(a) 在 ImageNet-1k(1.3M 图像)上,ViT 不如 ResNet(数据不足);(b) 在 JFT-300M(3 亿图像)上,ViT 超越 ResNet(且更大模型更好);(c) 交叉点约在数千万到上亿样本。其他关键差异——(a) 感受野——CNN 逐层扩大(线性),ViT 从第一层就全局;(b) 计算复杂度——CNN O(HW·k²·C²)、ViT O(N²·d)(长序列更贵);(c) 与文本的统一性——ViT 的 token 化使其易于与 LLM 拼接(这是 VLM 用 ViT 而非 CNN 的关键工程原因)。注入先验的手段——(a) 卷积 stem(用卷积做 patch embedding);(b) Swin 的窗口注意力(局部窗口 + 移动窗口);(c) 相对位置编码;(d) 混合架构(CNN + 注意力)。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Mathematical Mechanism: 1. CNN Inductive Biases: For feature map $X in mathbb{R}^{H times W times C_{text{in}}}$ and convolutional filter $W in mathbb{R}^{K times K times C_{text{in}} times C_{text{out}}}$: $$Y(i, j, c) = sum_{u=-k}^k sum_{v=-k}^k sum_{c’=1}^{C_{text{in}}} X(i+u, j+v, c’) cdot W(u, v, c’, c)$$ (a) Locality: Output depends strictly on the $K times K$ neighborhood. (b) Translation Equivariance: Let $T_v$ be spatial translation operator. Then $text{Conv}(T_v(X)) = T_v(text{Conv}(X))$. These structural constraints restrict hypothesis space $mathcal{H}_{text{CNN}} subset mathcal{H}_{text{ViT}}$. 2. Vision Transformer Formulation: Self-attention over all patch tokens: $$text{Attention}(Q, K, V) = text{Softmax}left( frac{Q K^T}{sqrt{d}} right) V, quad Q, K, V in mathbb{R}^{N times d}$$ At layer 1, any patch can attend directly to any distant patch across the entire image. Spatial topology is learned entirely through added positional encodings $E_{text{pos}}$. 3. Generalization & Sample Complexity: By PAC-learning theory, constrained hypothesis space $mathcal{H}_{text{CNN}}$ requires fewer training samples $m_{text{CNN}} = mathcal{O}(frac{text{VC}(mathcal{H}_{text{CNN}})}{epsilon})$ to avoid overfitting on small datasets. Conversely, unconstrained $mathcal{H}_{text{ViT}}$ exhibits lower approximation error on massive corpora: $$inf_{f in mathcal{H}_{text{ViT}}} mathcal{R}(f) < inf_{g in mathcal{H}_{text{CNN}}} mathcal{R}(g)$$ resulting in higher saturation limits when data volume $D to infty$.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
深度剖析与工程权衡:① ‘先验强度 vs 数据需求’是核心权衡——强先验(CNN)样本效率高但上限受限;弱先验(ViT)需数据但上限高。这解释了’ViT 需大数据、CNN 适合小数据’。② ‘VLM 用 ViT 的工程原因’——ViT 的 token 表示便于与文本 token 拼接(统一处理);CNN 的特征图需额外的’转 token’步骤。故现代 VLM 几乎都用 ViT 系(CLIP-ViT、SigLIP)。③ ‘自监督预训练改变了格局’——用自监督(MAE、DINO、CLIP)在大规模无标注数据上预训练 ViT,可大幅降低对标注数据的依赖;这是’ViT 数据饥渴’的主要解法(见自监督视觉预训练题)。④ ‘混合架构的实践’——ConvNeXt(用卷积模仿 Transformer 设计)、Swin(窗口注意力)在’样本效率 vs 灵活性’间折中;实践中’合适先验 + 足够数据’常优于纯架构之争。⑤ ‘小数据场景’——医疗/工业等小数据领域,CNN 或’预训练 ViT + 微调’更实用(从零训练 ViT 不现实)。⑥ 面试要点——被问’CNN vs ViT’,应给出’归纳偏置(局部性/平移等变 vs 全局/弱先验)→ 样本效率 → 数据交叉点‘的框架,并说明’VLM 用 ViT 是工程考虑(token 统一)‘与’自监督预训练降低了数据需求‘;这是 CV 基础类问题的核心。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Deep Dive & Engineering Trade-offs: ① The Data Scaling Law Shift: On ImageNet-1k (1.3M images) without pre-training or heavy regularization, ResNet-50 easily beats ViT-Base. When scaled to JFT-300M or LAION-5B with self-supervised or contrastive pre-training, ViT scales with power-law efficiency without saturating, while CNN performance plateaus. ② Multimodal Alignment Compatibility: ViTs output a discrete sequence of $N$ token vectors $z_1, dots, z_N in mathbb{R}^D$, perfectly matching the tokenized sequence abstraction of causal language models. Concatenating visual tokens with text tokens in Vision-Language Models (VLMs) is seamless with ViT, whereas CNN feature maps require spatial flattening, pooling, or custom cross-attention bridges. ③ Inference Compute Efficiency: While ViT features flexible receptive fields, its global attention scales as $mathcal{O}(N^2)$ with image resolution. CNNs scale linearly $mathcal{O}(HW cdot K^2)$ with resolution, making CNNs (or ConvNeXt) computationally lighter for high-resolution dense prediction tasks unless windowed attention is used. ④ Modern Hybrid Convergence: Modern architectures bridge this divide: ConvNeXt modernizes CNNs using $7 times 7$ depthwise convolutions and inverted bottlenecks, while Swin Transformer injects locality and hierarchical pyramid pooling into Transformers. ⑤ Interview Strategy: Formulate the convolution locality/equivariance equation, contrast it with global dot-product attention, cite the ImageNet vs JFT scaling crossover, and explain why ViTs dominate multimodal LLM backbones.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 在小数据上从零训练 ViT(欠拟合)
- ⚠️ 忽略 VLM 选 ViT 的工程动机(token 统一)
English Pitfalls:
– Attempting to train a vanilla ViT from scratch on small custom datasets (< 50K images) without heavy augmentations or pre-trained checkpoints
– Assuming CNNs cannot achieve competitive performance; modernized CNNs (ConvNeXt-V2) achieve comparable scaling when pre-trained with masked autoencoding
– Overlooking that ViTs possess zero innate spatial understanding and rely entirely on positional embeddings to distinguish top from bottom
六、高频深度面试追问与预测 (Follow-Up Questions)
- ViT 需要多少数据才能超越 CNN?
- Why does translation equivariance in CNNs fail to hold strictly under pooling layers and boundary padding artifacts?
- 如何给 ViT 注入视觉先验?
- How does ConvNeXt adapt Transformer design principles (patchify stem, $7times7$ depthwise convolutions, GeLU, LayerNorm) back into convolutional networks?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
Vision Transformer (ViT) 架构与图像 Patch 线性投影机制(Vision Transformers (ViT) & Patch Projection Mechanics) - 🗺️ 知识图谱模块:
深度学习架构导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。