Tag: architecture
-
多模态对齐:CLIP 双塔对比学习、InfoNCE 损失、Zero-Shot 迁移与 SigLIP 原理解构
> **核心摘要**:多模态表征对齐是连接视觉、文本、语音等异构数据的桥梁。OpenAI 提出的 **CLIP (Contrastive Language-Image Pre-Training)** 摒弃了传统的单模态分类标签,通过海量“图像-文本”对 (4 亿 pairs) 进行大规模双向对比学习,将视觉特征与自然语
-
离线强化学习与模仿学习:分布偏移、行为克隆、CQL、IQL 与 RLHF/DPO 全景全解
> **核心摘要**:离线强化学习 (Offline RL) 旨在从**固定、预收集的数据集** $mathcal{D} = {(s, a, r, s’)}$ 中学习策略,且**不再与环境交互**。其核心障碍是**分布偏移 (Distribution Shift)**:学到的策略会访问与数据分布不一致的状态-动作
-
AIE Agent Systems in Production: Orchestration Patterns, Context Budgeting, Reliability & Observability
> **Core Executive Summary**: An agent is a loop — goal, observation, policy, action, memory — but a production agent is a *guarded* loop. This guide covers the
-
MLE Coding & Algo Prep: Zero-to-One ML Operators in Pure Numpy
> **核心摘要**:Exhaustive technical deep dive into Pure Numpy handwritten ML operators for live coding interviews.
-
RS Paper Deep Dive Framework: Articulating Novelty & Research Vision
> **Executive Summary**: In Research Scientist (RS) interviews, interviewers evaluate a candidate’s **independent research taste, long-term technical vision, an
-
Cluster Scheduling & Ray: K8s Scheduling Pipeline, Raylet Architecture, Object Store, Autoscaling & GPU Scheduling Full Guide
> **Core Executive Summary**: Every ML platform asks one question: given $N$ nodes of CPU/GPU/memory, how do we place $M$ tasks so resources are utilized, users
-
Deep Learning Foundations: Activations Evolution (GELU/SwiGLU), Loss Function Taxonomy (CE/KL/Huber/InfoNCE/ArcFace) & Autograd Backprop Guide
> **Summary**: Non-linear activation functions, loss functions, and backpropagation form the mathematical pillars of deep learning. This exhaustive guide covers
-
Preference Alignment: RLHF 3-Stage, PPO Clipped Loss, DPO Math Derivation, GRPO & PRM/ORPO
> **Core Executive Summary**: While Pre-training and Supervised Fine-Tuning (SFT) instill strong language modeling and instruction following in LLMs, models may
-
Tokenizer & Decoding Strategies: BPE, WordPiece, SentencePiece, Temperature, Top-k/p, Min-p, Gumbel-Max, Repetition Penalty & Sequence Packing
> **Core Executive Summary**: The text processing pipeline of Large Language Models (LLMs) spans three stages: **Front-end Tokenization**, **Mid-end Autoregress
-
Generalization Theory: Inductive Bias, Double Descent & PAC Learning Paradigms
> **Core Executive Summary**: Why do 70B+ parameter LLMs generalize exceptionally without severe overfitting? **Generalization Theory** explains this modern AI