所属模块:
M2 · 经典机器学习 (Classical Machine Learning)| 专题分类:特征工程 (Feature Engineering)| 难度等级:Medium
一、核心一句话结论 (One-Sentence Summary)
滞后、滚动统计、周期性编码、时间差;必须保证只用过去信息(滚动窗口右对齐)。
Construct lag features, rolling window statistics, cyclical encodings, and elapsed durations while strictly ensuring right-aligned windows to avoid future data leakage.
二、核心考点要义 (Key Insights)
- 📌 sin/cos 编码周期避免边界跳变
- 📌 时间序列 CV 必须按时间切分
English Insights:
– Cyclical transformation: encode hour/day/month via sine and cosine to preserve periodicity
– Rolling window statistics: ensure windows are right-aligned using strictly historical data
– Temporal validation: use TimeSeriesSplit or forward-chaining, never random shuffle CV
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$text{rolling mean}: bar x_{t-1-w:t-1} (text{不含} t)$$
四类时间特征:① 时间戳分解——年/月/日/时/星期/是否节假日;注意星期与月份的周期性编码(用 sin(2πk/7)、cos(2πk/7) 而非 0–6 的整数,避免’周日与周一距离最大’的假象);② 滞后特征(lag)——过去 k 期的值(y_{t−1}, y_{t−2}…),捕捉自相关;③ 滚动统计——过去 w 期的均值/标准差/最大最小/趋势斜率,捕捉局部动态;④ 时间差特征——距上次事件的时间、事件频率、累积计数。泄漏的防范:核心原则是在时刻 t 只能用 t 之前(不含 t)的信息。具体做法:滚动窗口用 x[t-w:t](右开区间,不含 t)而非 x[t-w+1:t+1](含 t,泄漏);预测目标若是由 t 时刻信息决定的(如 t 时刻的转化),则特征必须严格早于该时点。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Core temporal feature categories: ① Cyclical Encoding: For periodic features (e.g., hour $h in [0, 23]$ or weekday $w in [0, 6]$), avoid linear integer encodings where 23 and 0 appear maximally distant. Map to 2D coordinates: $x_{sin} = sinleft(frac{2pi t}{T}right)$, $x_{cos} = cosleft(frac{2pi t}{T}right)$, preserving circular Euclidean distance. ② Lag Features: $x_{t – k}$ capturing autoregressive dependencies. ③ Rolling Aggregations: $mu_{t, W} = frac{1}{W} sum_{i=1}^W x_{t-i}$, strictly excluding $x_t$ if predicting target at time $t$.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
实践要点:① 最常见的泄漏写法——df.rolling(w).mean() 后再 shift(-1) 或直接对齐到当前行;正确做法是 df.shift(1).rolling(w).mean()(先滞后再滚动)或明确用右开区间。② 时间序列 CV——必须用前向链式 CV(训练集只含验证集之前的时间),随机 K 折会用到未来信息导致严重乐观偏差;同理,特征标准化/分箱等预处理也必须只在训练时段拟合。③ 周期性编码的选择——sin/cos 编码保留周期循环性;对非周期的时间(如’距项目开始的天数’)用原始数值即可。④ 节假日与事件——节假日、促销日、系统变更等外生事件对时间序列影响巨大,应显式建模(二值标记或事件类别特征)。⑤ 多序列与面板数据——若数据含多个实体(用户/商品),滚动特征需按实体分组计算(groupby 后再 rolling),且注意实体的生命周期(新实体无历史);这是实践中极易出错的点。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Production safeguards against leakage: ① Right-Aligned Windows: Rolling statistics must shift backward by at least the inference prediction horizon (e.g., if predicting 24 hours ahead, lags must start at $t-24$). ② Feature Arrival Delays: Account for logging latency in production pipelines (features might only become available hours after timestamp). ③ Cross-Validation Architecture: Strictly apply forward-chaining (expanding window or rolling window splits); never use standard randomized K-fold.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 用含当前时刻的滚动窗口(数据泄漏)
- ⚠️ 对时间序列用随机 K 折做交叉验证
English Pitfalls:
– Using future information in rolling window aggregations (e.g., centering the window at timestamp $t$)
– Evaluating time-series models using random K-fold cross-validation, leading to overly optimistic benchmark scores
六、高频深度面试追问与预测 (Follow-Up Questions)
- 为什么用 sin/cos 编码小时?
- Why is sine/cosine encoding necessary for hour-of-day features in distance-based and linear models?
- 滚动窗口常见的泄漏写法是什么?
- How do you handle production logging delays when designing real-time lag features?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
特征工程实战:Target Encoding、组合特征与特征离散化(Feature Engineering: Target Encoding & Feature Stores) - 🗺️ 知识图谱模块:
机器学习工程师高频考点导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。