AirSOTA
Air School of Thoughts AtoZAirSOTA 知识矩阵:聚合大模型算法架构、科学育儿情境成长、加州地产考牌实战与全球数字化商业出海的权威专栏。
TalentMe · AI 学习与系统架构
工业级 AI 算法核心 69 题、前沿大模型系统架构演进与北美技术面试全流程备考深度长文。
【AI 核心深度 M2-001】写出线性回归的闭式解,并说明它成立的前提。(Formulate the Closed-Form Normal Equation for Linear Regression and Its Necessary Preconditions)深度数理推导与工程落地解析
最小二乘解 w=(XᵀX)⁻¹Xᵀy,前提是 XᵀX 可逆(无完全共线性)。
【AI 核心深度 M2-002】列举线性回归的四大假设,并说明违反后果。(Enumerate the Four Classical Assumptions of Linear Regression and the Consequences of Their Violation)深度数理推导与工程落地解析
线性、误差独立、同方差、误差正态(用于推断)。
【AI 核心深度 M2-003】解释多重共线性,它如何影响系数估计与显著性检验。(Explain Multicollinearity, Its Impact on Coefficient Stability and Significance Tests, and Diagnostic Thresholds)深度数理推导与工程落地解析
特征高度相关 → XᵀX 接近奇异 → 系数方差爆炸、符号不稳定,但预测仍可能准。
【AI 核心深度 M2-004】岭回归为什么能缓解共线性?写出它的解。(Explain Why Ridge Regression Mitigates Multicollinearity and Derive Its Closed-Form Solution)深度数理推导与工程落地解析
在 XᵀX 上加 λI 使其可逆,同时收缩系数、稳定方差。
【AI 核心深度 M2-005】推导线性回归的最小二乘解,并说明其几何意义。(Derive the Least Squares Solution for Linear Regression and Explain Its Orthogonal Projection Geometry)深度数理推导与工程落地解析
对残差平方和求导置零得正规方程;几何上是把 y 正交投影到 X 的列空间。
【AI 核心深度 M2-006】写出逻辑回归的模型形式与损失函数。(Formulate Logistic Regression, Sigmoid Activation, and Binary Cross-Entropy Loss)深度数理推导与工程落地解析
sigmoid 输出概率,用交叉熵(对数似然)训练;是 GLM 中 Bernoulli + logit 链接。
【AI 核心深度 M2-007】解释 odds 与 log-odds,以及系数如何解释。(Explain Odds, Log-Odds (Logit), and the Exact Multiplicative Interpretation of Logistic Regression Coefficients)深度数理推导与工程落地解析
odds=p/(1-p),logit 是 log-odds;系数表示特征每增 1 单位,log-odds 变化 β。
【AI 核心深度 M2-008】什么是广义线性模型(GLM)?它由哪三部分组成。(Define Generalized Linear Models (GLMs) and Detail Their Three Foundational Components)深度数理推导与工程落地解析
GLM = 随机成分(指数族分布)+ 系统成分(线性预测子)+ 链接函数。
【AI 核心深度 M2-009】逻辑回归与单层神经网络是什么关系?(Explain the Structural Relationship Between Logistic Regression and a Single-Layer Neural Network)深度数理推导与工程落地解析
完全等价:单层线性 + sigmoid 输出,只是视角与优化方式不同。
【AI 核心深度 M2-010】如何处理逻辑回归中的完全分离(separation)问题?(Explain Complete Separation (Hauck-Donner Effect) in Logistic Regression and Standard Industrial Solutions)深度数理推导与工程落地解析
若某特征能完美分开两类,系数会趋向无穷、不收敛;需加正则或贝叶斯先验。
【AI 核心深度 M2-011】比较 L1、L2 与 ElasticNet 正则化的几何与效果差异。(Compare the Geometric and Functional Differences Between L1, L2, and ElasticNet Regularization)深度数理推导与工程落地解析
L1 稀疏(特征选择),L2 收缩(抗共线),ElasticNet 兼顾两者。
【AI 核心深度 M1-073】解释 FP32 主权重(master weights)与它在混合精度训练中的作用是什么。(Explain FP32 Master Weights and Their Essential Role in Mixed Precision Training (FP16/BF16))深度数理推导与工程落地解析
保留一份 FP32 参数副本用于更新,前向反向用低精度;防止小更新被舍入吞掉。
【AI 核心深度 M1-074】解释 Fisher 信息量与 Cramér-Rao 下界。(Explain Fisher Information and the Cramér-Rao Lower Bound (CRLB))深度数理推导与工程落地解析
Fisher 信息是对数似然关于参数的曲率(二阶导期望);CRB 给出无偏估计方差的下界。
【AI 核心深度 M1-075】什么是收缩估计(shrinkage)与 James-Stein 现象?(Explain Shrinkage Estimators and the James-Stein Phenomenon)深度数理推导与工程落地解析
把估计向先验/均值收缩以降低方差;James-Stein 表明在高维中收缩估计可优于无偏估计。
【AI 核心深度 M1-076】解释 t 检验与 z 检验的选择依据。(Explain the Decision Criteria Between Two-Sample t-Test and z-Test)深度数理推导与工程落地解析
z 检验用已知总体方差(或大样本近似);t 检验用样本方差(小样本、方差未知)。