Tag: module-m1
-
【AI 核心深度 M1-045】什么是序贯检验与 peeking 问题?如何正确做序贯实验。(Explain Sequential Testing and the Peeking Problem, and Detail Correct Continuous Monitoring Frameworks)深度数理推导与工程落地解析
反复查看并随时停止会极大抬高假阳性;需用 alpha-spending 或 always-valid 方法。
-
【AI 核心深度 M1-046】解释置信区间的频率派含义,以及它与 p-value 的关系。(Explain the Frequentist Interpretation of Confidence Intervals and Their Duality with Hypothesis Testing)深度数理推导与工程落地解析
95% CI 是’重复抽样下有 95% 的区间覆盖真值’;若 CI 不含 0,则对应双侧检验在 α=0.05 下显著。
-
【AI 核心深度 M1-048】什么是 bootstrap 的三种区间构造方法?(Compare the Three Standard Bootstrap Interval Construction Methods (Percentile, Normal, and BCa Intervals))深度数理推导与工程落地解析
百分位法、基本法(reverse percentile)、BCa(偏差校正加速)。
-
【AI 核心深度 M1-049】bootstrap 与置换检验(permutation test)分别适用于什么场景?(Contrast Bootstrap vs. Permutation Tests and Detail Their Respective Valid Assumptions and Use Cases)深度数理推导与工程落地解析
bootstrap 估计统计量的抽样分布(可给 CI);置换检验在 H0 下构造零分布(给 p 值)。
-
【AI 核心深度 M1-050】如何用 bootstrap 做 A/B 测试的显著性判断?有什么坑。(Detail How to Apply Bootstrap to Evaluate Statistical Significance in A/B Testing and Identify Key Pitfalls)深度数理推导与工程落地解析
对两组分别重采样,计算差值分布的分位数;坑:重尾、极小样本、多重指标、重复抽样需固定随机种子。
-
【AI 核心深度 M1-051】描述一个标准 A/B 实验的完整流程。(Describe the End-to-End Operational Lifecycle of a Standard Industrial A/B Experiment)深度数理推导与工程落地解析
定假设与指标 → 功效分析与样本量 → 随机分流 → 运行固定时长 → 校验分流 → 统计检验 → 决策与复盘。