Tag: foundations
-
MoE 混合专家模型与 DeepSeek MLA/MTP/mHC 架构解构:Top-k 门控、无辅助损失均衡、低秩潜注意力与 KAN 剖析
> **核心摘要**:随着模型参数量迈向万亿关口,稠密 (Dense) 模型前向计算开销急剧飙升。**混合专家模型 (Mixture-of-Experts, MoE)** 通过将 FFN 全连接层替换为多个独立的专家网络,并由**门控路由 (Gating Router)** 对每个 Token 仅动态激活前 $k$ 个
-
决策树与集成学习:CART、GBDT 负梯度拟合、XGBoost 二阶展开与 LightGBM 极客全解
> **核心摘要**:树模型与集成学习是表格数据 (Tabular Data) 领域的无冕之王。本指南系统梳理从单棵决策树的特征选择规则 (ID3 / C4.5 / CART) 到集成学习范式 (Bagging vs Boosting),深入剖析 GBDT 的负梯度拟合、XGBoost 的二阶泰勒展开与正则化叶子权重推
-
AI 数理基础全景:贝叶斯推断、香农信息熵、交叉熵与 KL 散度非对称证明
> **核心摘要**:概率论与信息论是整个人工智能、机器学习以及深度学习的核心数学支撑。从概率分布的**贝叶斯推断 (Bayesian Inference)** 到量化信息不确定性的**香农信息熵 (Entropy)**,再到深度学习损失函数的基石——**交叉熵 (Cross-Entropy)** 与 **KL 散度
-
智能体 RL 与推理搜索全景:MCTS 蒙特卡洛树搜索、PRM 过程监督与 RLVR 可验证奖励
> **核心摘要**:随着大语言模型迈向 **Agentic 自主智能体** 与 **System 2 慢思考** 阶段,传统单步静态输出已被长链轨迹规划 (Trajectory Planning)、多步工具调用与试错反思所取代。**智能体强化学习 (Agentic RL)** 将环境反馈与决策树搜索结合,形成了以 *
-
推荐系统工业级架构设计:召回-精排-重排三阶段、双塔模型与离在线一致性 Feature Store
> **核心摘要**:推荐系统是电商(淘宝/Amazon)、短视频(抖音/TikTok)以及信息流(小红书/Pinterest)的核心商业引擎。面对千万级 Item 与亿级 User,任何单一模型都无法在 **50ms 延迟 SLA** 内对全量候选打分,因此工业界采用**漏斗式多阶段架构:召回 (Retrieval)
-
RAG Pipeline: From Naive RAG to Advanced RAG Architecture, Hybrid Search, RRF & Cross-Encoder
> **Core Executive Summary**: Large Language Models suffer from knowledge cutoffs and hallucinations. **RAG (Retrieval-Augmented Generation)** connects LLMs to
-
KV Cache Management: Exact Bounds Derivation, vLLM PagedAttention & Prefix Caching
> **Core Executive Summary**: Autoregressive LLM generation requires caching key-value states to eliminate $O(N^2)$ recomputation. However, **KV Cache** imposes
-
Graph Neural Networks (GNN) Taxonomy: Graph Laplacian, Message Passing (MPNN), GCN, GraphSAGE, GAT & Edge Feature Guide
> **Summary**: Representation learning on non-Euclidean graph-structured data is fundamental to modern recommender systems and molecular modeling. This 100% exh
-
Classical NLP Tasks: NER, Text Classification, seq2seq Translation & NLI Entailment
> **Core Executive Summary**: Classical NLP established the foundations of text sequence modeling prior to large language models. From **NER (Named Entity Recog
-
Linear & Logistic Regression: Mathematical Derivations, Log-Odds, MLE, VIF & Bias-Variance Full Guide
> **Summary**: This comprehensive guide systematically covers the complete mathematical framework for Linear and Logistic Regression. We detail the 5 classical