🌐 RS Core Cheatsheet: Top 30 Papers Breakdown & Deep RL
Executive Summary: Technical interviews for Research Scientist (RS) roles evaluate first-principles mathematical rigor, analytical loss derivations, generalization bounds, and model failure modes. This cheatsheet deconstructs the RS core knowledge map: the 4-step paper breakdown framework, a structured matrix of the Top 30 milestone papers, analytical derivation of DPO implicit reward substitution, PPO clipped objective dynamics, and Pure Python implementations.
💡 Interactive Mermaid Architecture
graph TD
subgraph A["1. SOTA Paper Deconstruction Framework"]
A1["Step 1: Background & Problem Formulation"]
A2["Step 2: Core Bottlenecks of Prior Works"]
A3["Step 3: Key Mathematical Novelties"]
A4["Step 4: Decisive Ablations & Failure Modes"]
A1 --> A2 --> A3 --> A4
end
subgraph B["2. Deep RL & Alignment Evolution"]
B1["PPO: Clipped Surrogate Objective (First-order Trust Region Approximation)"]
B2["DPO: Closed-Form Optimal Policy (Implicit Reward Substitution)"]
B3["GRPO: Group Relative Advantage (Critic-free Pure RL Reasoning)"]
B1 --> B2 --> B3
end
A --> B
Chapter 1: The RS Core Mindset & 4-Step Paper Breakdown
RS interviews evaluate whether you can reason like the author of a pioneering breakthrough.
1. Problem Formulation: Concrete mathematical objective and scope
2. Physical Bottleneck: Fundamental scaling limits of prior paradigms
3. Algorithmic Novelty: Mathematical transformation and loss redesign
4. Decisive Ablations: Knockout experiments and out-of-distribution failure modes
Chapter 2: Pure Python PPO Clipped Loss Operator
import numpy as np
def pure_python_ppo_clipped_loss(ratios: np.ndarray, advantages: np.ndarray, clip_eps: float = 0.2) -> float:
surr1 = ratios * advantages
surr2 = np.clip(ratios, 1.0 - clip_eps, 1.0 + clip_eps) * advantages
return float(-np.mean(np.minimum(surr1, surr2)))
if __name__ == "__main__":
r = np.array([0.9, 1.15, 1.3])
adv = np.array([0.5, -0.4, 0.8])
print("✅ PPO Clipped Loss:", round(pure_python_ppo_clipped_loss(r, adv), 4))
Chapter 3: DPO Closed-Form Optimal Policy & Implicit Reward Derivation
Consider the KL-regularized RLHF objective:
$$max_{pi} mathbb{E}{x sim mathcal{D}, y sim pi(y|x)} [r(x, y)] – beta mathbb{D}(y|x))$$}}(pi(y|x) parallel pi_{text{ref}
Expanding the KL divergence:
$$max_{pi} sum_y pi(y|x) left( r(x, y) – beta log frac{pi(y|x)}{pi_{text{ref}}(y|x)} right) = max_{pi} -beta sum_y pi(y|x) log left( frac{pi(y|x)}{frac{1}{Z(x)} pi_{text{ref}}(y|x) exp(frac{1}{beta} r(x, y))} right) + beta log Z(x)$$
where partition function $Z(x) = sum_y pi_{text{ref}}(y|x) expleft( frac{1}{beta} r(x, y) right)$.
This is equivalent to minimizing the KL divergence to a Gibbs distribution. The global optimum is achieved when the distributions match, yielding the closed-form optimal policy:
$$pi^*(y|x) = frac{1}{Z(x)} pi_{text{ref}}(y|x) expleft( frac{1}{beta} r(x, y) right)$$
Taking logarithms and rearranging terms:
$$r(x, y) = beta log frac{pi^*(y|x)}{pi_{text{ref}}(y|x)} + beta log Z(x)$$
Substituting this implicit reward expression into the Bradley-Terry preference model $P(y_w succ y_l mid x) = sigma(r(x, y_w) – r(x, y_l))$, the partition constant $beta log Z(x)$ cancels out completely:
$$mathcal{L}{text{DPO}}(theta) = -mathbb{E} right) right]$$}} left[ log sigma left( beta log frac{pi_theta(y_w mid x)}{pi_{text{ref}}(y_w mid x)} – beta log frac{pi_theta(y_l mid x)}{pi_{text{ref}}(y_l mid x)
def pure_python_dpo_loss(
pi_yw: float, pi_yl: float,
ref_yw: float, ref_yl: float,
beta: float = 0.1
) -> float:
log_ratio_w = np.log(pi_yw) - np.log(ref_yw)
log_ratio_l = np.log(pi_yl) - np.log(ref_yl)
implicit_logit = beta * (log_ratio_w - log_ratio_l)
prob_win = 1.0 / (1.0 + np.exp(-implicit_logit))
return float(-np.log(prob_win))
if __name__ == "__main__":
print("✅ DPO Loss:", round(pure_python_dpo_loss(0.8, 0.2, 0.4, 0.4, beta=0.1), 4))
Chapter 4: Milestone Papers Taxonomy (Top 30 Works)
- Architecture: Attention Is All You Need (2017), FlashAttention-1/2 (2022/2023), DeepSeek-V3 MLA (2024).
- Scaling: Chinchilla (2022), DeepSeek-R1 (2025), Kaplan Scaling Laws (2020).
- Alignment: InstructGPT (2022), DPO (2023), GRPO (2024).
- Diffusion & Generation: DDPM (2020), Flow Matching (2023), DiT / Sora (2023/2024).
🧠 深入探索 TalentMe 全景技术图谱与备考路线
本文选自 TalentMe AI 技术专栏与高维职业罗盘。支持双模态 Obsidian 本地私域同步、艾宾浩斯智能复习与 IDE 内嵌 AI 导师模拟面试。