所属模块:
M6 · 多模态与生成模型 (Multimodal & Generative Models)| 专题分类:视频 / 3D / 音频 (Video, 3D & Audio Generative Models)| 难度等级:Medium
一、核心一句话结论 (One-Sentence Summary)
NeRF 用 MLP 表示隐式辐射场(体积渲染,慢);3DGS 用显式高斯基元(光栅化,快且可实时)。
NeRF models 3D scenes as continuous implicit neural volumetric functions optimized via ray marching, while 3D Gaussian Splatting represents scenes using explicit anisotropic 3D Gaussians rendered via fast tile-based rasterization.
二、核心考点要义 (Key Insights)
- 📌 NeRF:隐式(MLP 查询密度与颜色)+ 体积渲染(慢)
- 📌 3DGS:显式(大量 3D 高斯基元)+ 光栅化(快)
- 📌 3DGS 可实时渲染、训练更快;NeRF 更紧凑但渲染慢
English Insights:
– NeRF representation (Mildenhall et al.): encodes continuous scene geometry and view-dependent color using a coordinate MLP $,F_theta(x, d) to (sigma, c),$, rendering via numerical volume integration along camera rays
– 3D Gaussian Splatting (Kerbl et al.): models scenes using millions of explicit anisotropic 3D Gaussians defined by position, covariance matrix, opacity, and spherical harmonics
– Rendering speed breakthrough: NeRF requires evaluating hundreds of MLP forward passes per pixel (seconds per frame), while 3DGS achieves real-time $100+,$ FPS rendering via GPU tile-based rasterization
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$text{NeRF}: F_theta(x,d)to(c,sigma);qquad text{3DGS}: {(mu_i,Sigma_i,c_i,alpha_i)} text{rasterized}$$
数学机理:NeRF(Neural Radiance Fields,Mildenhall 等 2020)——(1) 隐式表示——用一个 MLP F_θ(x,d)→(c,σ) 表示场景:输入 3D 位置 x 与视角方向 d,输出颜色 c 与体密度 σ。(2) 体积渲染——对每条光线沿路径采样多个点,用体渲染公式积分得到像素颜色:C=∫T(t)σ(t)c(t)dt(离散化为求和)。(3) 训练——用多视角图像做监督(最小化渲染图与真实图的差异);通过可微渲染反向传播到 MLP。(4) 为什么慢——(a) 每条光线需采样数十到数百个点,每点都要一次 MLP 前向(计算量大);(b) 渲染一张图需数十万到数百万次 MLP 前向;(c) 故 NeRF 渲染需数秒到数十秒(不可实时)。(5) 改进——(a) 位置编码(把 x 映射到高频基函数,使 MLP 能表示细节);(b) Instant-NGP(用哈希网格替代 MLP,大幅加速训练与渲染);(c) 分层采样(在’有物体的区域’多采样)。3D Gaussian Splatting(3DGS,Kerbl 等 2023)——(1) 显式表示——用大量 3D 高斯基元(数万到数百万个)表示场景:每个基元有位置 μ、协方差 Σ(形状/朝向)、颜色 c(含球谐系数以支持视角相关颜色)、不透明度 α。(2) 渲染——把这些高斯基元投影到屏幕空间(2D 高斯)并光栅化(按深度排序 + alpha 混合);无需逐光线采样 + MLP 前向。(3) 为什么快——(a) 光栅化比’体积渲染 + MLP’快得多(GPU 光栅化硬件友好);(b) 可实现实时渲染(>30 FPS);(c) 训练更快(分钟级 vs NeRF 的数小时)。(4) 优点——实时、训练快、可编辑(直接改基元);(5) 缺点——(a) 存储大(数百万基元需大量显存);(b) 显式表示(不是’连续场’,故在稀疏视角下泛化差);(c) 对’反射/透明’等复杂材质仍困难。对比总结——(a) 表示:隐式(MLP)vs 显式(基元);(b) 渲染:体积渲染(慢)vs 光栅化(快);(c) 速度:NeRF 秒级 vs 3DGS 实时;(d) 存储:NeRF 小(一个 MLP)vs 3DGS 大(数百万基元);(e) 训练:NeRF 小时级 vs 3DGS 分钟级;(f) 编辑性:3DGS 更易编辑(改基元);(g) 泛化:NeRF 类(隐式)在稀疏视角下更好。与生成模型的结合——(a) 3D 生成(如 DreamFusion 用 2D 扩散 + NeRF/3DGS 优化);(b) 3DGS 的生成(如用扩散生成 3DGS 参数);(c) 前馈式重建(如 LRM、VGGT 直接从图像预测 3D)。评估——(a) 渲染质量(PSNR/SSIM/LPIPS);(b) 渲染速度(FPS);(c) 存储大小;(d) 训练时间。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Mathematical Mechanism: 1. NeRF Volumetric Ray Marching: For camera ray $r(t) = o + t d$, rendered pixel color $C(r)$ integrates continuous volume density $sigma(r(t))$ and color $c(r(t), d)$ between near bound $t_n$ and far bound $t_f$: $$C(r) = int_{t_n}^{t_f} T(t) sigma(r(t)) c(r(t), d) dt, quad T(t) = expleft( -int_{t_n}^t sigma(r(s)) ds right)$$ Approximated numerically via stratified quadrature sampling across $K$ discrete points: $$hat{C}(r) = sum_{i=1}^K T_i (1 – exp(-sigma_i delta_i)) c_i, quad T_i = prod_{j=1}^{i-1} exp(-sigma_j delta_j)$$ 2. 3D Gaussian Splatting Formulation: A 3D Gaussian is parameterized by mean position $mu in mathbb{R}^3$ and 3D covariance matrix $Sigma in mathbb{R}^{3 times 3}$: $$G(x) = expleft( -frac{1}{2} (x – mu)^T Sigma^{-1} (x – mu) right)$$ To ensure positive semi-definiteness during gradient descent, covariance is decomposed into scaling vector $s in mathbb{R}^3$ and rotation quaternion $q in mathbb{R}^4$: $$Sigma = R(q) S(s) S(s)^T R(q)^T$$ 3. 2D Projective Splatting (Zwicker et al.): Given viewing transformation $W$ and projective Jacobian $J$, the 2D projected screen covariance is: $$Sigma’ = J W Sigma W^T J^T in mathbb{R}^{2 times 2}$$ Pixels are colored via tile-based alpha-blending of sorted Gaussians: $$C = sum_{i in mathcal{N}} c_i alpha_i prod_{j=1}^{i-1} (1 – alpha_j), quad alpha_i = o_i cdot G_{2text{D}}(p – mu’_i)$$
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
深度剖析与工程权衡:① ‘显式 vs 隐式’是核心区分——它决定了渲染速度、存储与泛化;面试中能给出这一对比是深度理解的标志。② ‘3DGS 可实时’是它的杀手级优势——使 3D 内容可用于交互式应用(VR/游戏);这是 NeRF 难以达到的。③ ‘存储代价’是 3DGS 的代价——数百万基元需大量显存/磁盘;故有压缩方法(如量化、剪枝)。④ ‘稀疏视角下 NeRF 类更好’——隐式表示有’先验平滑性’(因为 MLP 连续),故在少视角时泛化更好;3DGS 是显式的(需足够视角)。⑤ ‘与生成模型的结合是当前热点’——2D 扩散提供’语义先验’,3D 表示提供’几何一致性’;故 (a) 用扩散监督 NeRF/3DGS 优化(SDS)、(b) 用前馈模型直接预测 3D。⑥ 面试要点——被问’NeRF vs 3DGS’,应给出’隐式 MLP + 体积渲染(慢)vs 显式高斯 + 光栅化(实时)‘与’存储/泛化/编辑性的对比‘;能指出’3DGS 实时但存储大、NeRF 泛化好但慢’是深度理解的标志。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Deep Dive & Engineering Trade-offs: ① The Paradigm Shift from Implicit to Explicit: NeRF was revolutionary in 2020 for eliminating discrete polygonal meshes and achieving photorealistic novel view synthesis. However, querying a deep MLP 128 times per ray for a $1920 times 1080$ frame ($2 times 10^6$ rays) requires hundreds of millions of neural network evaluations per image, capping rendering speed at 0.1-1 FPS. 3D Gaussian Splatting eliminates neural networks at inference entirely: rendering is pure GPU rasterization (sorting 2D Gaussians and executing hardware alpha blending), reaching 100-200 FPS on consumer hardware at equal or higher visual quality. ② Memory Footprint and Storage: (a) NeRF: Compact representation; the entire scene is compressed into a $5text{–}50,text{MB}$ MLP weight checkpoint. (b) 3DGS: Storing 2-5 million 3D Gaussians (each storing position, rotation, scale, opacity, and 16 spherical harmonic coefficients) consumes $500,text{MB}$ to $1.5,text{GB}$ of RAM/disk storage. Compression techniques (vector quantization of spherical harmonics, Gaussian pruning) reduce footprint by $10times$. ③ Adaptive Density Control in 3DGS: 3DGS periodically splits large Gaussians in under-reconstructed regions and clones small Gaussians in high-frequency detail regions, while pruning transparent Gaussians ($o_i < epsilon$). ⑤ Interview Strategy: Contrast implicit volumetric MLP rendering against explicit Gaussian primitive rasterization, write NeRF’s volume rendering quadrature equation, derive 3DGS covariance projection $Sigma’ = J W Sigma W^T J^T$, and compare rendering speed vs memory footprint.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 以为 NeRF 可实时渲染(需数十万次 MLP 前向)
- ⚠️ 忽略 3DGS 的存储代价
English Pitfalls:
– Attempting real-time interactive rendering with standard vanilla NeRF; vanilla NeRF requires seconds to render a single frame
– Optimizing covariance matrix $Sigma$ directly without scale-rotation decomposition ($R S S^T R^T$), causing non-positive-definite covariance collapse
– Underestimating the storage footprint of raw uncompressed 3D Gaussian splat files ($1text{GB}+$) in web deployment pipelines
六、高频深度面试追问与预测 (Follow-Up Questions)
- NeRF 为什么慢?
- Why is the scale-rotation decomposition $Sigma = R S S^T R^T$ mathematically necessary when optimizing 3D Gaussian Splatting?
- 3DGS 的’高斯基元’是什么?
- How does tile-based GPU rasterization enable 3D Gaussian Splatting to render complex scenes at over 100 FPS?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
时空视频扩散架构、3D 高斯泼溅 (3DGS) 与语音音频生成模型(Spatiotemporal Video Diffusion, 3DGS & Audio Generation) - 🗺️ 知识图谱模块:
多模态与扩散模型导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。