Tag: foundations
-
Optimizers & Training Engineering Taxonomy: SGD, Momentum, AdamW Decoupled Weight Decay, Xavier/Kaiming Initialization & Gradient Checkpointing Guide
> **Summary**: Optimizers and training engineering bridge the gap between network architecture design and physical GPU memory limits. This 100% exhaustive guide
-
LLM Quantization & Model Compression: INT8/INT4 Mapping, SmoothQuant Outliers, GPTQ Hessian & AWQ/Distillation
> **Core Executive Summary**: As LLM parameter counts scale into tens to hundreds of billions, FP16/BF16 VRAM consumption and memory bandwidth become severe lat
-
Probabilistic Graphical Models: Naive Bayes, HMM Viterbi & Linear-Chain CRF Guide
> **Summary**: Probabilistic Graphical Models (PGM) combine graph theory and probability theory. This guide covers Naive Bayes conditional independence, HMM Vit
-
Speech & Audio Processing: Whisper Architecture, Log-Mel Spectrogram & Audio-LLM
> **Core Executive Summary**: Speech is the most natural medium for human interaction. Traditional Speech Recognition (ASR) relied on complex acoustic and langu
-
Model-Based RL & Planning: World Models, Dyna, MPC, MuZero & Dreamer
> **Core Executive Summary**: Model-based reinforcement learning (MBRL) equips the agent with an internal world model — a learned approximation of the transitio
-
TalentMe AI/ML/LLM Full Knowledge Taxonomy & Architecture Graph
> **Overview & Vision**: Modern Artificial Intelligence and Large Language Models have evolved into a massive, interdisciplinary, and mathematically rigorous en
-
工具调用与 Function Calling 全景:Toolformer 自主插入、JSON Schema 规范与沙箱安全执行
> **核心摘要**:大语言模型(LLM)虽然具备强大的文本生成能力,但无法实时查询当前天气、无法直接进行精确的大数字浮点运算,也无法直接执行代码。**Tool Use (工具调用)** 与 **Function Calling** 突破了 LLM 的能力边界,使其能够通过结构化 JSON 规范与外部 API、数据库以
-
AI 安全与隐私全景:Prompt 注入攻击、Guardrails 防御、差分隐私与联邦学习
> **核心摘要**:大语言模型(LLM)的开放交互特性带来了前所未有的安全挑战。**Prompt 注入攻击** 能够绕过系统设定劫持模型行为,**PII 泄露** 可能引发严重的合规危机。通过在输入输出端部署 **Guardrails (安全护栏)**,并在模型微调阶段引入 **差分隐私 (DP-SGD)** 与 *
-
优化器与训练工程全景:SGD、Momentum、AdamW 解耦权重衰减、Xavier/Kaiming 初始化推导、梯度累积与重计算 (Checkpointing) 极客指南
> **核心摘要**:优化器与训练工程是连接神经网络结构设计与显存硬件物理限制的坚实桥梁。从自适应学习率优化器的演进(SGD $to$ Momentum $to$ RMSprop $to$ Adam $to$ AdamW)、打破对称性迷局的权重初始化数理推导(Xavier/Glorot 方差守恒与 Kaimin
-
大模型量化与模型压缩全景:INT8/INT4 映射、SmoothQuant 异常值平滑、GPTQ 二阶 Hessian 优化与 AWQ/知识蒸馏剖析
> **核心摘要**:随着大语言模型 (LLM) 参数量达到百亿至千亿级,全精度 FP16/BF16 模型的显存占用与访存带宽成为实时低延迟推理的致命瓶颈。**模型量化 (Quantization)** 技术通过将连续高精度浮点数映射为低精度整数(如 INT8、INT4),在显存占用缩减 50%~75% 的同时实现显着