Category: AI & 机器学习 (AI & Machine Learning)
-
GPU Hardware Architecture: SM, Tensor Cores, HBM Bandwidth & Roofline Model
> **Core Executive Summary**: AI LLM performance relies directly on underlying GPU hardware physics. Modern GPUs like NVIDIA H100/A100 feature massively paralle
-
Deep Learning Debugging & Competition Engineering Taxonomy: 4-Step Debugging Framework, Single Batch Overfitting, Gradient Check & Grad-CAM Guide
> **Summary**: Debugging deep learning models is notoriously challenging because bad code often runs without crashing while silently degrading performance. This
-
Open & Commercial SOTA LLM Evolution: From BERT/GPT-4 to LLaMA-3, Qwen-3, Gemma-4 & Kimi-K2
> **Core Executive Summary**: Since the Transformer paper, Large Language Models (LLMs) evolved from unidirectional/bidirectional encoders (BERT/GPT-1/2) to lar
-
Unsupervised Clustering & KNN: K-Means++, DBSCAN, GMM-EM & KD-Tree Guide
> **Summary**: Clustering and nearest-neighbor methods form the backbone of pattern recognition and representation analysis. This guide explores coordinate desc
-
Optimization & Matrix Calculus: Lagrange Multipliers, KKT Conditions, SVD & Convergence Geometry
> **Core Executive Summary**: Every machine learning model—from SVM geometric margin maximization to Transformer gradient descent updates—is fundamentally an **
-
World Models & JEPA: Yann LeCun’s Non-Generative Prediction, I-JEPA/V-JEPA & Embodied AI (VLA)
> **Core Executive Summary**: Turing Award winner Yann LeCun proposed **JEPA (Joint Embedding Predictive Architecture)**, advocating abandoning pixel-level reco
-
Multimodal Generative System Design: Image/Video Generation & GPU Scaling
> **Core Executive Summary**: Multimodal generative workloads — text-to-image, text-to-video, and vision-language understanding — are the most compute-hungry an
-
AIE 实战指南:企业级 SFT、LoRA 微调与 DPO/RLHF 偏好对齐落地
> **核心摘要**:大模型微调与对齐是 AI 工程师(AIE)将开源基座模型落地为垂类业务专家的核心技能。在工业实践中,微调面临三大挑战:训练数据长短不一导致 Padding 算力巨大浪费、显存不足导致全量微调代价高昂、以及强化学习对齐训练极度不稳定。本指南系统剖析 Data Packing 算法、Loss Mask
-
MLE 数据与特征工程:质量管线、编码策略、不平衡处理、特征选择与漂移检测全景
> **核心摘要**:模型的性能在训练开始之前就已注定——它由数据质量与特征工程决定。本指南覆盖 MLE 数据管线的完整闭环:数据质量四大威胁(漏标、噪声、重复、异常值)与清洗管线、缺失值填充策略对比、类别特征编码(Label / One-Hot / Target / OOF / Embedding)与防泄漏、数值特征
-
Agent 设计模式全景:ReAct 循环、Reflexion 自我反思、Plan-and-Execute 与 LangGraph 图工程
> **核心摘要**:大语言模型(LLM)不仅能作为无状态的问答工具,更能进化为具备自主决策能力的 **AI Agent (智能体)**。Agent 通过四大支柱——**Brain (大脑推理)**、**Planning (规划解构)**、**Memory (长短期记忆)** 和 **Tools (工具调用)**,能够