AirSOTA
Air School of Thoughts AtoZAirSOTA 知识矩阵:聚合大模型算法架构、科学育儿情境成长、加州地产考牌实战与全球数字化商业出海的权威专栏。
TalentMe · AI 学习与系统架构
工业级 AI 算法核心 69 题、前沿大模型系统架构演进与北美技术面试全流程备考深度长文。
大模型偏好对齐全景:RLHF 3 阶段、PPO 截断损失、DPO 隐式奖励代换、GRPO 与 PRM/ORPO 深度剖析
> **核心摘要**:预训练与 Supervised Fine-Tuning (SFT) 赋予了大语言模型 (LLM) 强大的语言建模与指令遵循能力,但模型依然可能生成有害、偏见或不符合人类期望的回复(即 **Alignment Tax 现象**)。人类偏好对齐技术通过引入人类或 AI 反馈,引导模型向**有用性 (H
Tokenizer 分词器与 LLM 采样解码全景:BPE、WordPiece、SentencePiece、Temperature、Top-k/p、Min-p、Gumbel-Max、Penalty 与 Sequence Packing 打包优化
> **核心摘要**:大语言模型 (LLM) 的文本处理管道由**前端文本离散化 (Tokenization)**、**中端自回归 Logit 预测**与**后端概率采样解码 (Decoding & Sampling)** 三大核心阶段构成。本指南系统剖析 BPE、WordPiece、Unigram、SentenceP
泛化理论全景:归纳偏置 (Inductive Bias)、Double Descent 双重下降与 PAC 学习范式
> **核心摘要**:为什么深度学习模型(如 700 亿参数大模型)拥有远超训练样本数的参数量,却不会发生严重的过拟合,反而展现出惊人的泛化能力?**泛化理论 (Generalization Theory)** 解释了这一现代 AI 奇迹。通过 **Inductive Bias (归纳偏置)** 注入先验架构约束,借助
扩散模型全景:DDPM 数理推导、Latent Diffusion (LDM)、DiT 架构与 GPT-4o Native 生成
> **核心摘要**:生成式 AI 的两大支柱分别是自回归模型 (Autoregressive LLMs) 与 **扩散模型 (Diffusion Models)**。扩散模型借鉴了非平衡态热力学 (Non-equilibrium Thermodynamics) 原理,通过向数据添加高斯噪声(前向过程)并学习一步步恢复
业界经典 System Case Studies:Pinterest 视觉搜索与 Netflix 推荐系统
> **核心摘要**:学习 System Design 的最高境界是研读业界顶级科技巨头的真实架构。本指南全量解构两个经典工业案例——**Pinterest**(视觉搜索与推荐:图像嵌入、PinSage 式图神经网络表征、HNSW 近似最近邻检索、混合检索、多模态表征)与 **Netflix**(流媒体推荐:显式 +
AIE Core Cheatsheet: SFT, LoRA, RAG & Agent Interview Map
> **Executive Summary**: The AI / LLM Systems Engineer (AIE) role spans model fine-tuning, retrieval engineering, agentic orchestration, and high-throughput ser
MLE Core Cheatsheet: High-Frequency Q&A, Competitions & Pinterest
> **Core Executive Summary**: The MLE interview core is a closed loop of recurring topics — **regularization & bias-variance, overfitting diagnosis, feature eng
Writing_Guide
本规范适用于 `content/tech/` 下所有双语主题文件(`*.zh.md` / `*.en.md`)。目标:**面试导向** —— 每个知识点不仅要有公式,还要有”能直接说出口的面试回答”。
Distributed Training Parallelism: TP, PP, DP & DeepSpeed ZeRO 1/2/3
> **Core Executive Summary**: Single GPU VRAM cannot host 100B+ LLM training parameters, gradients, and optimizer states (a 70B FP16 model requires 1.12TB train
Vision Architectures Evolution: 2D Conv, Receptive Field Calculus, Depthwise Separable Conv, ResNet Identity Mapping & Vision Transformer (ViT) Guide
> **Summary**: Computer vision architectures evolved from handcrafted local inductive biases (CNNs) to data-driven global self-attention (ViT). This 100% exhaus
LLM Hallucination & Factuality: Taxonomies, FActScore, RAGAS, SAFE & Context Extension (PI/NTK/YaRN)
> **Core Executive Summary**: LLMs often generate plausible-sounding but unfactual or logically contradictory text, known as **Hallucination**. Hallucinations r
Transformer Architecture Breakdown: Self-Attention, MHA/GQA/MQA, RoPE & FlashAttention 1/2/3 Operator Fusion
> **Core Executive Summary**: Since its introduction in 2017, the Transformer architecture has fundamentally reshaped artificial intelligence, serving as the un
Linear Algebra Core for AI: Vector Spaces, Four Subspaces, EVD/SVD, Projection & Least Squares, Jacobian/Hessian
> **Core Executive Summary**: Linear algebra is the substrate of machine learning and deep learning: every tensor is a matrix, every layer is a matrix multiplic
Vision-Language Models (VLM): ViT, Projectors, LLaVA 2-Stage & DeepSeek-Janus Pro
> **Core Executive Summary**: Vision-Language Models (VLMs) empower LLMs to perceive visual scenes. Rather than training multimodal models from scratch, VLMs ut
Production LLM RAG & Agent System Design: Multi-Tenancy, SSE & High Availability
> **Core Executive Summary**: Moving RAG knowledge bases and agentic systems from demo to enterprise production is a distributed-systems problem as much as an M