所属模块:
M5 · NLP 与大语言模型 (NLP & Large Language Models)| 专题分类:Tokenization (Tokenization (BPE / WordPiece / Unigram))| 难度等级:Hard
一、核心一句话结论 (One-Sentence Summary)
为新语言/领域追加 token 并初始化其 embedding,再继续训练;注意初始化方式、学习率与对原能力的保护。
Vocabulary expansion adds domain- or language-specific subwords to a pre-trained model by resizing the embedding and LM head matrices, requiring careful weight initialization and targeted continual pre-training to prevent catastrophic forgetting.
二、核心考点要义 (Key Insights)
- 📌 追加 token → 扩展 embedding 与输出层 → 初始化新行 → 继续预训练
- 📌 新 token 的 embedding 常用旧 token 的均值/同方差采样初始化
- 📌 注意:原 token 的分词分布会变化(需重新适配)
English Insights:
– Motivation: adapting an existing pre-trained LLM (e.g., LLaMA-1/2 with poor Chinese support) to a new language or domain by adding $10text{k}text{–}30text{k}$ new subword tokens
– Matrix resizing: expands embedding matrix $W_E in mathbb{R}^{V times d}$ to $W_E’ in mathbb{R}^{(V + K) times d}$ and adjusts output LM head $W_{text{head}}$
– Initialization strategy: initialize new token embeddings by averaging the embeddings of their constituent old subwords rather than using random Gaussian initialization
– Continual pre-training: freeze base Transformer layers initially to align new embeddings, then unfreeze all layers with small learning rates
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$e_{text{new}}simmathcal{N}(mu_{text{old}},sigma_{text{old}});qquad text{then continue pretraining with lower lr}$$
数学机理:动机——原 tokenizer 对目标语言/领域效率低(fertility 高);扩展词表可 (a) 降低 token 数(省推理成本)、(b) 提升上下文有效容量、(c) 让模型更高效地学习该语言/领域。做法:(1) 训练新 tokenizer 或增量合并——在目标语料上训练一个 tokenizer,或把新子词并入原词表(保持原 token 的 id 不变,新 token 追加在后);(2) 扩展参数矩阵——把输入 embedding 矩阵 E∈ℝ^{V×d} 扩展为 E’∈ℝ^{(V+K)×d},输出层同理(若 tied 则一起扩展);(3) 初始化新行——常用 (a) 均值初始化(新 token 的向量 = 原词表 embedding 的均值)、(b) 同分布采样(从原 embedding 的均值与方差采样),避免用随机初始化(会引入分布外的大梯度);(4) 继续预训练(continued pretraining)——用目标语言/领域语料继续训练,学习率较小(如原预训练 lr 的 1/10),步数视数据量而定;可只训练 embedding 与输出层(冻结其余)或全参数微调。注意点:(a) 原 token 的分布变化——新词表改变了分词方式,原语言的文本也会被切成不同的 token 序列(虽然原 token id 未变),故模型需’重新适应’新分词分布;(b) 原能力的保护——继续预训练可能导致灾难性遗忘,故常混合原语言数据(如 10%~30%)以保持原能力;(c) 输出层与 tied embedding——若 tied,扩展需同步(否则输入输出不一致);(d) 位置编码与上下文——扩展词表不改变最大长度,但’每 token 字节数’变化会影响有效容量。效果——对低资源语言/新领域,词表扩展通常比’不改词表直接继续训练’更高效(序列更短、学习更快)。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Mathematical Mechanism: 1. Subword Compositional Initialization: Let new token $t_{text{new}}$ be added to vocabulary. Decompose $t_{text{new}}$ into its sub-tokens under the old tokenizer $mathcal{T}_{text{old}}$: $t_{text{new}} = [u_1, u_2, dots, u_m]$, where $u_i in mathcal{V}_{text{old}}$. Initialize the new embedding vector $e(t_{text{new}})$ as the mean of its constituent old embeddings: $$e(t_{text{new}}) = frac{1}{m} sum_{i=1}^m e_{text{old}}(u_i)$$ This places new tokens in the semantic neighborhood of their components, preventing massive gradient shock at step 0. 2. Output Head Initialization: Similarly, initialize the output projection vector $w_{text{head}}(t_{text{new}})$ from the mean of constituent output vectors: $$w_{text{head}}(t_{text{new}}) = frac{1}{m} sum_{i=1}^m w_{text{head}}(u_i)$$
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
深度剖析与工程权衡:① ‘新 token 的初始化’是关键细节——用随机初始化会让新 token 在训练初期产生异常大的梯度(因为其 embedding 远离训练分布),可能破坏模型;均值/同分布初始化使新 token 从’合理的起点’开始学习。② 词表扩展 vs 从头训练 tokenizer——扩展保留原模型能力(推荐);从头换 tokenizer 等于’重训整个模型’(不现实)。故实践中’扩展’是主流。③ 与 LoRA/PEFT 的组合——若只想加新语言能力,可’扩展词表 + 继续训练 embedding + LoRA 微调其余层’;这比全参数继续预训练便宜得多。④ 对原语言的性能影响——扩展后原语言的分词会变化(新词表可能改变合并规则);若新词表是在’原语料 + 新语料’上重训的,可能损害原语言的压缩率;故常采用’增量合并‘(保留原合并规则、只追加新 token)以最小化影响。⑤ 评估方式——需分别测 (a) 目标语言的 token 效率(fertility/压缩率)、(b) 目标语言的任务指标、(c) 原语言的任务指标(确保未遗忘);三者缺一不可。⑥ 面试要点——被问’如何让模型支持新语言’,应给出’扩展词表(增量合并 + 均值初始化)+ 继续预训练(小 lr + 混合原语料)+ 评估双向指标‘的流程,并说明’均值初始化’与’灾难性遗忘防护’两个关键点;这是’工程落地’类问题的高分回答。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Deep Dive & Engineering Trade-offs: ① Random vs Mean Initialization: Random Gaussian initialization causes severe gradient spikes on the first training step, as uninitialized embeddings produce chaotic output logits that propagate massive gradients backward through all layers. Mean subword initialization stabilizes loss from step 1. ② Staged Training Pipeline (Chinese-LLaMA Paradigm): – Stage 1 (Embedding Warmup): Freeze all Transformer attention and MLP layers; train only the newly added embedding and head weights ($W_E’, W_{text{head}}’$) on domain text for a few billion tokens. – Stage 2 (Full Alignment): Unfreeze all model parameters, training with a small learning rate (e.g., $10%$ of original pre-training LR) and a 15-20% replay mixture of original pre-training data to prevent catastrophic forgetting. ③ Tokenizer Training Mixture: When training the new expansion tokenizer, train it on a mixture of target language text and original English text to preserve high English compression ratios. ④ Modern Relevance: In modern frontier LLMs (LLaMA-3, Qwen-2), native vocabularies ($128text{k}+$) already provide rich multilingual coverage, reducing the necessity of manual post-hoc vocabulary expansion. ⑤ Interview Strategy: Describe the matrix resizing mechanics, derive the subword averaging initialization formula, and explain the two-stage training regimen (freeze base $to$ unfreeze with replay).
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 用随机初始化新 token 的 embedding
- ⚠️ 继续预训练时不混合原语料(导致遗忘)
English Pitfalls:
– Initializing new token embeddings with random Gaussian noise (causes massive gradient shock and destroys pre-trained representations)
– Continual pre-training without original-language replay data (leads to severe catastrophic forgetting of general reasoning)
– Unfreezing all Transformer layers immediately before new embeddings have aligned their representation magnitudes
六、高频深度面试追问与预测 (Follow-Up Questions)
- 为什么新 token 要用均值初始化?
- Why is freeze-base warmup necessary during the first stage of vocabulary expansion?
- 词表扩展对原能力有什么风险?
- How does Chinese-LLaMA merge its expanded vocabulary with the original LLaMA tokenizer without collision?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
分词算法与原理:BPE 字节对编码、WordPiece、Unigram 与多语言分词(Tokenization Algorithms: BPE, WordPiece & Multilingual Vocab) - 🗺️ 知识图谱模块:
大语言模型全景图谱
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。