【AI 核心深度 M5-002】比较 BPE、WordPiece 与 Unigram。(Comparison of BPE, WordPiece, and Unigram Tokenization)深度数理推导与工程落地解析

所属模块:M5 · NLP 与大语言模型 (NLP & Large Language Models) | 专题分类:Tokenization (Tokenization (BPE / WordPiece / Unigram)) | 难度等级:Easy

一、核心一句话结论 (One-Sentence Summary)

BPE 按频次合并;WordPiece 按’似然增益’合并(BERT);Unigram 从大词表出发按概率剪枝(T5/SentencePiece)。

ADVERTISEMENT · 赞助推荐

BPE merges subwords based on raw co-occurrence frequency, WordPiece merges subwords based on mutual information (maximum likelihood), and Unigram prunes a large initial candidate pool based on sentence-level likelihood loss.

二、核心考点要义 (Key Insights)

  • 📌 BPE:频次驱动的合并(自底向上)
  • 📌 WordPiece:似然比驱动的合并(偏好’共同出现’的组合)
  • 📌 Unigram:概率模型 + 剪枝(自顶向下),支持多种分词采样

English Insights:
– BPE: bottom-up; greedy merge maximizing co-occurrence frequency $text{count}(t_i, t_j)$; deterministic and simple (GPT, LLaMA, RoBERTa)
– WordPiece: bottom-up; greedy merge maximizing training corpus unigram likelihood improvement $frac{text{count}(t_i, t_j)}{text{count}(t_i) text{count}(t_j)}$ (BERT)
– Unigram: top-down; starts with a massive vocabulary and iteratively prunes tokens that cause the smallest drop in corpus unigram language model likelihood (SentencePiece, T5, LLaMA-3)

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$text{WordPiece score}(a,b)=frac{mathrm{count}(ab)}{mathrm{count}(a)mathrm{count}(b)};qquad text{Unigram}: mathcal{L}=sumlog p(text{subword})$$

数学机理:三种算法的差异在’合并/剪枝的判据’与’方向’。(1) BPE——自底向上,合并频次最高的相邻对:merge*=argmax count(a,b)。问题——只考虑绝对频次,可能合并’虽然频繁但语义上不构成整体’的组合(如 de 在 de+f 与 ide 中含义不同)。(2) WordPiece(BERT 用)——自底向上,但合并判据是似然比(点互信息式):score(a,b)=count(ab)/(count(a)·count(b))。直觉——它偏好’单独出现少、但一起出现多’的组合(即’共同出现’的强关联),而非单纯高频;这使合并更’有意义’(如 ing 作为后缀的关联性强)。(3) Unigram(SentencePiece 的默认,T5/ALBERT/LLaMA 用)——自顶向下:先构造一个很大的候选子词集合(如所有高频子串),用一元语言模型(每个子词独立概率,句子的概率是子词概率之积)在语料上最大化似然训练概率,然后剪枝掉’删除后损失增加最小’的子词,逐步缩到目标词表。关键优势:(a) 概率模型使同一文本可有多种分词(按概率采样),支持subword regularization(训练时随机采样分词,提升鲁棒性)与 BPE-dropout;(b) 自顶向下的剪枝可全局权衡(不像 BPE 的贪心)。共同点——都产生子词词表、都解决 OOV、都被 SentencePiece/tiktoken 等库实现。实践——BPE 最简单最快(GPT 系用 tiktoken 的 BPE/BBPE);WordPiece 是 BERT 的遗产;Unigram 在 T5/LLaMA 系常用,且支持采样增强。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Mathematical Mechanism: 1. BPE Merge Criterion: $$text{Score}(t_i, t_j) = text{Freq}(t_i, t_j)$$ Merges the pair that occurs most frequently in the raw text corpus. 2. WordPiece Merge Criterion: Evaluates the increase in corpus log-likelihood when merging $t_i$ and $t_j$: $$text{Score}(t_i, t_j) = frac{text{Freq}(t_i, t_j)}{text{Freq}(t_i) times text{Freq}(t_j)}$$ This measures Pointwise Mutual Information (PMI). High-frequency individual tokens (like ‘the’) are penalized unless their combination appears far more frequently than independent chance. 3. Unigram Pruning Criterion: Assumes independent unigram probabilities $p(t) = frac{text{count}(t)}{N}$. Corpus log-likelihood is $mathcal{L} = sum_{x} log P(x) = sum_{x} log left( max_{S in text{Seg}(x)} prod_{t in S} p(t) right)$. For each candidate token $t$: compute the drop in likelihood $Delta mathcal{L}_t = mathcal{L}_{mathcal{V}} – mathcal{L}_{mathcal{V} setminus {t}}$. The bottom $p%$ (e.g., $10text{–}20%$) of tokens with the smallest $Delta mathcal{L}_t$ are pruned. Repeats until target size $V$ is reached.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

深度剖析与工程权衡:① ‘自底向上 vs 自顶向下’的本质差异——BPE/WordPiece 从字符开始合并(局部贪心、无法撤销);Unigram 从大集合开始剪枝(全局权衡、可重新评估)。这使 Unigram 在理论上更接近’最优词表’,但训练成本更高(需 EM 迭代)。② subword regularization 的价值——Unigram 的概率模型允许’同一句有多种分词’,训练时随机采样可 (a) 提升对分词差异的鲁棒性、(b) 起到数据增强作用;BPE 需专门的 BPE-dropout 才能实现类似效果。③ 词表共享与多语言——多语言模型(mBERT、XLM-R、mT5)需在多语言语料上联合训练 tokenizer;此时词表分配(每种语言分到多少 token)成为公平性问题(低资源语言被压缩得更厉害)。④ 确定性对可复现性的影响——BPE/WordPiece 的分词是确定性的(给定词表),便于缓存与复现;Unigram 若启用采样则非确定(训练时),推理时通常用 Viterbi 最优分词(确定性)。⑤ 与 SentencePiece 的关系——SentencePiece 把’预处理(空格处理)’与’分词’统一(把空格也视为普通符号 ▁),使分词可逆(可精确还原原文);这对多语言与代码很重要(避免’去空格/加空格’的不可逆处理)。⑥ 面试要点——被问’三种分词的区别’,应给出’合并判据(频次 / 似然比 / 概率剪枝)+ 方向(自底向上 / 自顶向下)‘,并说明’Unigram 支持 subword regularization’与’SentencePiece 保证可逆’;能指出’多语言词表分配公平性’是深度理解的标志。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

Deep Dive & Engineering Trade-offs: ① Deterministic vs Probabilistic Segmentation: BPE and WordPiece are bottom-up greedy algorithms yielding a single deterministic segmentation per word. Unigram naturally models a probability distribution over multiple valid segmentations for a sentence, enabling subword regularization (sampling different segmentations during training to boost robustness). ② Inference Complexity: BPE applies ordered merge rules ($O(L log |mathcal{V}|)$). Unigram uses the Viterbi algorithm over a dynamic programming trellis to find the most probable token sequence in $O(L cdot L_{max})$. ③ SentencePiece Library Dominance: Google’s SentencePiece library implements both BPE and Unigram natively in C++, treating whitespace as a regular character (`_` or ` `) to avoid language-dependent pre-tokenization. ④ Subword Regularization: Unigram sampling introduces data augmentation during training: the same sentence ‘unhappiness’ can be seen as `[‘un’, ‘happiness’]` or `[‘un’, ‘happy’, ‘ness’]`, dramatically reducing overfitting on rare words. ⑤ Interview Strategy: Contrast bottom-up merging (BPE/WordPiece) with top-down pruning (Unigram), compare the score metrics (Frequency vs PMI vs Likelihood Drop), and explain subword regularization.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 以为 WordPiece 与 BPE 只是命名不同(判据不同)
  • ⚠️ 忽略 Unigram 的采样增强能力

English Pitfalls:
– Believing WordPiece merges tokens based on raw frequency like BPE (WordPiece normalizes by individual token frequencies via PMI)
– Assuming Unigram builds its vocabulary bottom-up (Unigram is top-down pruning from an oversized initial candidate set)
– Overlooking that Unigram natively supports subword regularization via Viterbi trellis sampling

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 为什么 WordPiece 用似然比而不是频次?
  2. How does Unigram’s Viterbi segmentation find the globally optimal subword breakdown for a sentence?
  3. Unigram 的 EM 如何优化?
  4. Why does subword regularization improve downstream translation and classification accuracy?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:分词算法与原理:BPE 字节对编码、WordPiece、Unigram 与多语言分词 (Tokenization Algorithms: BPE, WordPiece & Multilingual Vocab)
  • 🗺️ 知识图谱模块:大语言模型全景图谱

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M5-002) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.