Tag: foundations
-
凸优化与矩阵求导全景:拉格朗日乘子法、KKT 条件、SVD 奇异值分解与梯度几何收敛
> **核心摘要**:每一个机器学习模型(无论是 SVM 的几何间隔最大化、PCA 的方差最大化,还是 Transformer 的梯度下降更新)本质上都是一个**约束或无约束数学优化问题**。**矩阵求导 (Matrix Calculus)**、**KKT 条件 (Karush-Kuhn-Tucker)** 和 **S
-
世界模型与 JEPA 全景:Yann LeCun 非生成式表征预测、I-JEPA / V-JEPA 与具身智能 (VLA) 落地
> **核心摘要**:图灵奖得主 Yann LeCun 提出了著名的 **JEPA (Joint Embedding Predictive Architecture)** 架构,主张放弃像素级生成(Pixel Reconstruction),转而在抽象表征空间 (Representation Space) 中预测世界的
-
多模态生成系统架构设计:图像/视频生成服务、模型切片与 GPU 动态扩缩容
> **核心摘要**:多模态生成负载(文生图、文生视频、视觉语言理解)是当今 ML 服务中计算量最大、延迟最高的场景,绝不能沿用经典同步 RPC 架构。单张 SDXL 图像需要在数十亿参数的 UNet/DiT 主干上迭代 25+ 步去噪(单卡 A100 需 2–4 秒),而一段 5 秒 720p 视频需要处理数百万时空
-
LLM-as-a-Judge Evaluation: Pointwise & Pairwise Paradigms, Bias Elimination & Cohen’s Kappa
> **Core Executive Summary**: Traditional metrics like BLEU and ROUGE fail to evaluate complex semantic quality. **LLM-as-a-Judge** uses strong LLMs (such as GP
-
High-Concurrency AI System Design: SSE Streaming, Semantic Cache & ML Runtimes
> **Core Executive Summary**: Traditional web servers handle millisecond HTTP requests. LLM serving involves multi-second streaming responses. **High-Concurrenc
-
Generative Adversarial Networks (GAN) Taxonomy: Minimax Game, JS Divergence Flaw, WGAN Earth Mover Distance & WGAN-GP Guide
> **Summary**: Generative Adversarial Networks (GANs) frame generative modeling as a two-player zero-sum game between a Generator and Discriminator. This 100% e
-
Mixture-of-Experts (MoE) & DeepSeek MLA/MTP/mHC Architecture: Top-k Routing, Aux-Loss-Free, KAN vs MLP
> **Core Executive Summary**: As model scales reach trillion-parameter frontiers, Dense forward FLOPs become unsustainable. **Mixture-of-Experts (MoE)** replace
-
Decision Trees & Ensemble Methods: CART, GBDT 2nd-Order Taylor & LightGBM Guide
> **Summary**: Tree-based ensemble methods represent the state of the art for tabular datasets. This guide explores decision tree splitting criteria (ID3 / C4.5
-
AI Math Foundations: Bayes Inference, Shannon Entropy, Cross-Entropy & KL Divergence
> **Core Executive Summary**: Probability theory and information theory form the mathematical backbone of artificial intelligence. From **Bayesian Inference** p
-
Agentic RL & Reasoning Search: MCTS, Process Supervision & RLVR
> **Core Executive Summary**: As LLMs evolve toward **Autonomous Agents** and **System 2 Slow-Thinking**, static single-pass generation gives way to trajectory