【AI 核心深度 M6-094】解释 NeRF 与 3D Gaussian Splatting 的差异。(Neural Radiance Fields (NeRF) vs 3D Gaussian Splatting: Continuous Volumetric vs Discrete Primitives)深度数理推导与工程落地解析

所属模块:M6 · 多模态与生成模型 (Multimodal & Generative Models) | 专题分类:视频 / 3D / 音频 (Video, 3D & Audio Generative Models) | 难度等级:Medium

一、核心一句话结论 (One-Sentence Summary)

NeRF 用 MLP 表示隐式辐射场(体积渲染,慢);3DGS 用显式高斯基元(光栅化,快且可实时)。

ADVERTISEMENT · 赞助推荐

NeRF models 3D scenes as continuous implicit neural volumetric functions optimized via ray marching, while 3D Gaussian Splatting represents scenes using explicit anisotropic 3D Gaussians rendered via fast tile-based rasterization.

二、核心考点要义 (Key Insights)

  • 📌 NeRF:隐式(MLP 查询密度与颜色)+ 体积渲染(慢)
  • 📌 3DGS:显式(大量 3D 高斯基元)+ 光栅化(快)
  • 📌 3DGS 可实时渲染、训练更快;NeRF 更紧凑但渲染慢

English Insights:
– NeRF representation (Mildenhall et al.): encodes continuous scene geometry and view-dependent color using a coordinate MLP $,F_theta(x, d) to (sigma, c),$, rendering via numerical volume integration along camera rays
– 3D Gaussian Splatting (Kerbl et al.): models scenes using millions of explicit anisotropic 3D Gaussians defined by position, covariance matrix, opacity, and spherical harmonics
– Rendering speed breakthrough: NeRF requires evaluating hundreds of MLP forward passes per pixel (seconds per frame), while 3DGS achieves real-time $100+,$ FPS rendering via GPU tile-based rasterization

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$text{NeRF}: F_theta(x,d)to(c,sigma);qquad text{3DGS}: {(mu_i,Sigma_i,c_i,alpha_i)} text{rasterized}$$

数学机理:NeRF(Neural Radiance Fields,Mildenhall 等 2020)——(1) 隐式表示——用一个 MLP F_θ(x,d)→(c,σ) 表示场景:输入 3D 位置 x 与视角方向 d,输出颜色 c 与体密度 σ。(2) 体积渲染——对每条光线沿路径采样多个点,用体渲染公式积分得到像素颜色:C=∫T(t)σ(t)c(t)dt(离散化为求和)。(3) 训练——用多视角图像做监督(最小化渲染图与真实图的差异);通过可微渲染反向传播到 MLP。(4) 为什么慢——(a) 每条光线需采样数十到数百个点,每点都要一次 MLP 前向(计算量大);(b) 渲染一张图需数十万到数百万次 MLP 前向;(c) 故 NeRF 渲染需数秒到数十秒(不可实时)。(5) 改进——(a) 位置编码(把 x 映射到高频基函数,使 MLP 能表示细节);(b) Instant-NGP(用哈希网格替代 MLP,大幅加速训练与渲染);(c) 分层采样(在’有物体的区域’多采样)。3D Gaussian Splatting(3DGS,Kerbl 等 2023)——(1) 显式表示——用大量 3D 高斯基元(数万到数百万个)表示场景:每个基元有位置 μ、协方差 Σ(形状/朝向)、颜色 c(含球谐系数以支持视角相关颜色)、不透明度 α。(2) 渲染——把这些高斯基元投影到屏幕空间(2D 高斯)并光栅化(按深度排序 + alpha 混合);无需逐光线采样 + MLP 前向。(3) 为什么快——(a) 光栅化比’体积渲染 + MLP’快得多(GPU 光栅化硬件友好);(b) 可实现实时渲染(>30 FPS);(c) 训练更快(分钟级 vs NeRF 的数小时)。(4) 优点——实时、训练快、可编辑(直接改基元);(5) 缺点——(a) 存储大(数百万基元需大量显存);(b) 显式表示(不是’连续场’,故在稀疏视角下泛化差);(c) 对’反射/透明’等复杂材质仍困难。对比总结——(a) 表示:隐式(MLP)vs 显式(基元);(b) 渲染:体积渲染(慢)vs 光栅化(快);(c) 速度:NeRF 秒级 vs 3DGS 实时;(d) 存储:NeRF 小(一个 MLP)vs 3DGS 大(数百万基元);(e) 训练:NeRF 小时级 vs 3DGS 分钟级;(f) 编辑性:3DGS 更易编辑(改基元);(g) 泛化:NeRF 类(隐式)在稀疏视角下更好。与生成模型的结合——(a) 3D 生成(如 DreamFusion 用 2D 扩散 + NeRF/3DGS 优化);(b) 3DGS 的生成(如用扩散生成 3DGS 参数);(c) 前馈式重建(如 LRM、VGGT 直接从图像预测 3D)。评估——(a) 渲染质量(PSNR/SSIM/LPIPS);(b) 渲染速度(FPS);(c) 存储大小;(d) 训练时间。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Mathematical Mechanism: 1. NeRF Volumetric Ray Marching: For camera ray $r(t) = o + t d$, rendered pixel color $C(r)$ integrates continuous volume density $sigma(r(t))$ and color $c(r(t), d)$ between near bound $t_n$ and far bound $t_f$: $$C(r) = int_{t_n}^{t_f} T(t) sigma(r(t)) c(r(t), d) dt, quad T(t) = expleft( -int_{t_n}^t sigma(r(s)) ds right)$$ Approximated numerically via stratified quadrature sampling across $K$ discrete points: $$hat{C}(r) = sum_{i=1}^K T_i (1 – exp(-sigma_i delta_i)) c_i, quad T_i = prod_{j=1}^{i-1} exp(-sigma_j delta_j)$$ 2. 3D Gaussian Splatting Formulation: A 3D Gaussian is parameterized by mean position $mu in mathbb{R}^3$ and 3D covariance matrix $Sigma in mathbb{R}^{3 times 3}$: $$G(x) = expleft( -frac{1}{2} (x – mu)^T Sigma^{-1} (x – mu) right)$$ To ensure positive semi-definiteness during gradient descent, covariance is decomposed into scaling vector $s in mathbb{R}^3$ and rotation quaternion $q in mathbb{R}^4$: $$Sigma = R(q) S(s) S(s)^T R(q)^T$$ 3. 2D Projective Splatting (Zwicker et al.): Given viewing transformation $W$ and projective Jacobian $J$, the 2D projected screen covariance is: $$Sigma’ = J W Sigma W^T J^T in mathbb{R}^{2 times 2}$$ Pixels are colored via tile-based alpha-blending of sorted Gaussians: $$C = sum_{i in mathcal{N}} c_i alpha_i prod_{j=1}^{i-1} (1 – alpha_j), quad alpha_i = o_i cdot G_{2text{D}}(p – mu’_i)$$

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

深度剖析与工程权衡:① ‘显式 vs 隐式’是核心区分——它决定了渲染速度、存储与泛化;面试中能给出这一对比是深度理解的标志。② ‘3DGS 可实时’是它的杀手级优势——使 3D 内容可用于交互式应用(VR/游戏);这是 NeRF 难以达到的。③ ‘存储代价’是 3DGS 的代价——数百万基元需大量显存/磁盘;故有压缩方法(如量化、剪枝)。④ ‘稀疏视角下 NeRF 类更好’——隐式表示有’先验平滑性’(因为 MLP 连续),故在少视角时泛化更好;3DGS 是显式的(需足够视角)。⑤ ‘与生成模型的结合是当前热点’——2D 扩散提供’语义先验’,3D 表示提供’几何一致性’;故 (a) 用扩散监督 NeRF/3DGS 优化(SDS)、(b) 用前馈模型直接预测 3D。⑥ 面试要点——被问’NeRF vs 3DGS’,应给出’隐式 MLP + 体积渲染(慢)vs 显式高斯 + 光栅化(实时)‘与’存储/泛化/编辑性的对比‘;能指出’3DGS 实时但存储大、NeRF 泛化好但慢’是深度理解的标志。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

Deep Dive & Engineering Trade-offs: ① The Paradigm Shift from Implicit to Explicit: NeRF was revolutionary in 2020 for eliminating discrete polygonal meshes and achieving photorealistic novel view synthesis. However, querying a deep MLP 128 times per ray for a $1920 times 1080$ frame ($2 times 10^6$ rays) requires hundreds of millions of neural network evaluations per image, capping rendering speed at 0.1-1 FPS. 3D Gaussian Splatting eliminates neural networks at inference entirely: rendering is pure GPU rasterization (sorting 2D Gaussians and executing hardware alpha blending), reaching 100-200 FPS on consumer hardware at equal or higher visual quality. ② Memory Footprint and Storage: (a) NeRF: Compact representation; the entire scene is compressed into a $5text{–}50,text{MB}$ MLP weight checkpoint. (b) 3DGS: Storing 2-5 million 3D Gaussians (each storing position, rotation, scale, opacity, and 16 spherical harmonic coefficients) consumes $500,text{MB}$ to $1.5,text{GB}$ of RAM/disk storage. Compression techniques (vector quantization of spherical harmonics, Gaussian pruning) reduce footprint by $10times$. ③ Adaptive Density Control in 3DGS: 3DGS periodically splits large Gaussians in under-reconstructed regions and clones small Gaussians in high-frequency detail regions, while pruning transparent Gaussians ($o_i < epsilon$). ⑤ Interview Strategy: Contrast implicit volumetric MLP rendering against explicit Gaussian primitive rasterization, write NeRF’s volume rendering quadrature equation, derive 3DGS covariance projection $Sigma’ = J W Sigma W^T J^T$, and compare rendering speed vs memory footprint.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 以为 NeRF 可实时渲染(需数十万次 MLP 前向)
  • ⚠️ 忽略 3DGS 的存储代价

English Pitfalls:
– Attempting real-time interactive rendering with standard vanilla NeRF; vanilla NeRF requires seconds to render a single frame
– Optimizing covariance matrix $Sigma$ directly without scale-rotation decomposition ($R S S^T R^T$), causing non-positive-definite covariance collapse
– Underestimating the storage footprint of raw uncompressed 3D Gaussian splat files ($1text{GB}+$) in web deployment pipelines

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. NeRF 为什么慢?
  2. Why is the scale-rotation decomposition $Sigma = R S S^T R^T$ mathematically necessary when optimizing 3D Gaussian Splatting?
  3. 3DGS 的’高斯基元’是什么?
  4. How does tile-based GPU rasterization enable 3D Gaussian Splatting to render complex scenes at over 100 FPS?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:时空视频扩散架构、3D 高斯泼溅 (3DGS) 与语音音频生成模型 (Spatiotemporal Video Diffusion, 3DGS & Audio Generation)
  • 🗺️ 知识图谱模块:多模态与扩散模型导图

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M6-094) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.