RS Core Cheatsheet: Top 30 Papers Breakdown & Deep RL

EN
This technical guide is also available in Chinese.


🌐 查看中文版本 / Read in Chinese →

🌐 RS Core Cheatsheet: Top 30 Papers Breakdown & Deep RL

Executive Summary: Technical interviews for Research Scientist (RS) roles evaluate first-principles mathematical rigor, analytical loss derivations, generalization bounds, and model failure modes. This cheatsheet deconstructs the RS core knowledge map: the 4-step paper breakdown framework, a structured matrix of the Top 30 milestone papers, analytical derivation of DPO implicit reward substitution, PPO clipped objective dynamics, and Pure Python implementations.


💡 Interactive Mermaid Architecture

graph TD
    subgraph A["1. SOTA Paper Deconstruction Framework"]
        A1["Step 1: Background & Problem Formulation"]
        A2["Step 2: Core Bottlenecks of Prior Works"]
        A3["Step 3: Key Mathematical Novelties"]
        A4["Step 4: Decisive Ablations & Failure Modes"]
        A1 --> A2 --> A3 --> A4
    end

    subgraph B["2. Deep RL & Alignment Evolution"]
        B1["PPO: Clipped Surrogate Objective (First-order Trust Region Approximation)"]
        B2["DPO: Closed-Form Optimal Policy (Implicit Reward Substitution)"]
        B3["GRPO: Group Relative Advantage (Critic-free Pure RL Reasoning)"]
        B1 --> B2 --> B3
    end

    A --> B

Chapter 1: The RS Core Mindset & 4-Step Paper Breakdown

RS interviews evaluate whether you can reason like the author of a pioneering breakthrough.

ADVERTISEMENT · 赞助推荐

1. Problem Formulation: Concrete mathematical objective and scope
2. Physical Bottleneck: Fundamental scaling limits of prior paradigms
3. Algorithmic Novelty: Mathematical transformation and loss redesign
4. Decisive Ablations: Knockout experiments and out-of-distribution failure modes

Chapter 2: Pure Python PPO Clipped Loss Operator

import numpy as np

def pure_python_ppo_clipped_loss(ratios: np.ndarray, advantages: np.ndarray, clip_eps: float = 0.2) -> float:
    surr1 = ratios * advantages
    surr2 = np.clip(ratios, 1.0 - clip_eps, 1.0 + clip_eps) * advantages
    return float(-np.mean(np.minimum(surr1, surr2)))

if __name__ == "__main__":
    r = np.array([0.9, 1.15, 1.3])
    adv = np.array([0.5, -0.4, 0.8])
    print("✅ PPO Clipped Loss:", round(pure_python_ppo_clipped_loss(r, adv), 4))

Chapter 3: DPO Closed-Form Optimal Policy & Implicit Reward Derivation

Consider the KL-regularized RLHF objective:
$$max_{pi} mathbb{E}{x sim mathcal{D}, y sim pi(y|x)} [r(x, y)] – beta mathbb{D}(y|x))$$}}(pi(y|x) parallel pi_{text{ref}

Expanding the KL divergence:
$$max_{pi} sum_y pi(y|x) left( r(x, y) – beta log frac{pi(y|x)}{pi_{text{ref}}(y|x)} right) = max_{pi} -beta sum_y pi(y|x) log left( frac{pi(y|x)}{frac{1}{Z(x)} pi_{text{ref}}(y|x) exp(frac{1}{beta} r(x, y))} right) + beta log Z(x)$$
where partition function $Z(x) = sum_y pi_{text{ref}}(y|x) expleft( frac{1}{beta} r(x, y) right)$.

This is equivalent to minimizing the KL divergence to a Gibbs distribution. The global optimum is achieved when the distributions match, yielding the closed-form optimal policy:
$$pi^*(y|x) = frac{1}{Z(x)} pi_{text{ref}}(y|x) expleft( frac{1}{beta} r(x, y) right)$$

Taking logarithms and rearranging terms:
$$r(x, y) = beta log frac{pi^*(y|x)}{pi_{text{ref}}(y|x)} + beta log Z(x)$$

Substituting this implicit reward expression into the Bradley-Terry preference model $P(y_w succ y_l mid x) = sigma(r(x, y_w) – r(x, y_l))$, the partition constant $beta log Z(x)$ cancels out completely:
$$mathcal{L}{text{DPO}}(theta) = -mathbb{E} right) right]$$}} left[ log sigma left( beta log frac{pi_theta(y_w mid x)}{pi_{text{ref}}(y_w mid x)} – beta log frac{pi_theta(y_l mid x)}{pi_{text{ref}}(y_l mid x)

def pure_python_dpo_loss(
    pi_yw: float, pi_yl: float,
    ref_yw: float, ref_yl: float,
    beta: float = 0.1
) -> float:
    log_ratio_w = np.log(pi_yw) - np.log(ref_yw)
    log_ratio_l = np.log(pi_yl) - np.log(ref_yl)
    implicit_logit = beta * (log_ratio_w - log_ratio_l)
    prob_win = 1.0 / (1.0 + np.exp(-implicit_logit))
    return float(-np.log(prob_win))

if __name__ == "__main__":
    print("✅ DPO Loss:", round(pure_python_dpo_loss(0.8, 0.2, 0.4, 0.4, beta=0.1), 4))

Chapter 4: Milestone Papers Taxonomy (Top 30 Works)

  • Architecture: Attention Is All You Need (2017), FlashAttention-1/2 (2022/2023), DeepSeek-V3 MLA (2024).
  • Scaling: Chinchilla (2022), DeepSeek-R1 (2025), Kaplan Scaling Laws (2020).
  • Alignment: InstructGPT (2022), DPO (2023), GRPO (2024).
  • Diffusion & Generation: DDPM (2020), Flow Matching (2023), DiT / Sora (2023/2024).

🧠 深入探索 TalentMe 全景技术图谱与备考路线

本文选自 TalentMe AI 技术专栏与高维职业罗盘。支持双模态 Obsidian 本地私域同步、艾宾浩斯智能复习与 IDE 内嵌 AI 导师模拟面试。

👉 访问 TalentMe 技术专栏 →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.