所属模块:
M6 · 多模态与生成模型 (Multimodal & Generative Models)| 专题分类:动态分辨率与视觉 token (Dynamic Resolution & Visual Tokens)| 难度等级:Medium
一、核心一句话结论 (One-Sentence Summary)
把高分辨率图切成多个固定尺寸的 tile(各自编码)+ 一个全局缩略图,兼顾细节与全局;代价是 tile 数与 token 数增长。
Image tiling partitions high-resolution inputs into discrete sub-image crops to preserve native pixel density, balancing fine-grained local perception against token budget inflation and inter-tile boundary fragmentation.
二、核心考点要义 (Key Insights)
- 📌 tile:切成多个固定尺寸子图,各自过 ViT(保留细节)
- 📌 全局缩略图:提供整图的全局视野(避免只见局部)
- 📌 tile 数 k 由原图尺寸决定 → token 数 ∝k
English Insights:
– Tiling mechanics: slices ultra-high-resolution images into $M times N$ grid patches of standard resolution (e.g., $336 times 336$ or $448 times 448$) matching the vision encoder’s pre-trained size
– Dual-stream overview: pairs local detailed tiles with a resized global overview crop to anchor macro scene understanding and inter-tile spatial continuity
– Boundary fragmentation trade-off: objects sliced across tile borders suffer from visual cleavage, requiring overlap margins or LLM cross-attention stitching
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$text{AnyRes}: {t_1,dots,t_k}+text{global thumbnail};qquad text{tokens}=ktimes N_{text{tile}}+N_{text{thumb}}$$
数学机理:tiling / AnyRes 的机制——(1) 切分——把高分辨率图切成 k 个固定尺寸的 tile(如每个 336×336 或 448×448);切分方式通常按宽高比选择网格(如 1×2、2×2、3×3),使每个 tile 接近方形。(2) 各自编码——每个 tile 独立过 ViT(+连接器),得到 N_tile 个 token;共 k×N_tile 个 token。(3) 全局缩略图——把整图缩放到一个固定尺寸(如 336×336)再编码,得到 N_thumb 个 token(约 N_tile),提供全局视野。为什么需要全局缩略图——因为 tile 各自独立编码,模型看不到 tile 之间的关系(不知道’这个 tile 在图中的位置’);全局缩略图提供’整图的样子’,帮助模型建立全局理解(如’这是文档的左上角’)。总 token 数——k×N_tile + N_thumb;故 token 数 ∝ k ∝ 图像面积(与动态分辨率类似,token 与尺寸成正比)。与’原生分辨率(不切分)’的对比——(a) tiling——保持 ViT 的固定输入尺寸(利于复用预训练权重),但需处理 tile 间关系与位置编码;(b) 原生分辨率——直接按原生尺寸切 patch(ViT 需支持可变输入,位置编码需 2D/M-RoPE)。tiling 的优点——(a) 复用固定尺寸的预训练 ViT(不需改视觉塔);(b) 实现简单(切图 + 批处理);(c) 可处理任意尺寸(通过选择 tile 网格)。tiling 的缺点——(a) tile 间关系丢失(需全局缩略图或位置编码弥补);(b) 边界切断(物体可能被切在 tile 边界);(c) token 数增长快(k 大时昂贵);(d) tile 内的相对位置需额外编码(否则模型不知道 tile 在图中的位置)。缓解——(a) tile 位置编码(给每个 tile 加一个’位置嵌入’标记其在网格中的位置);(b) 全局缩略图(提供全局视野);(c) token 压缩(对 tile token 做池化);(d) 重叠切分(避免边界切断,但 token 更多)。代表——(a) LLaVA-NeXT 的 AnyRes——用 1×1 到 3×3 的网格 + 全局图;(b) InternVL 的动态 tiling(自适应选择网格);(c) Qwen-VL 用原生分辨率(不用 tiling)。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Mathematical Mechanism: 1. Grid Partitioning Formulation: Let input image be $I in mathbb{R}^{H times W times 3}$ and standard encoder crop size be $S times S$. An adaptive grid $(m, n)$ is selected with constraint $m cdot n le K_{max}$: $$m = text{round}left(frac{H}{S}right), quad n = text{round}left(frac{W}{S}right)$$ The image is resized to $(m cdot S, n cdot S)$ using bicubic interpolation with antialiasing, and sliced into $m times n$ non-overlapping tiles: $$I_{text{tiles}} = { I_{i, j} mid i in {1, dots, m}, ; j in {1, dots, n} }, quad I_{i, j} in mathbb{R}^{S times S times 3}$$ 2. Global Context Stream: Concurrently, a downsampled global overview is generated: $$I_{text{global}} = text{Resize}(I, (S, S))$$ 3. Feature Sequence Assembly: Vision encoder $f_v$ processes all $m cdot n + 1$ crops independently. Tokens are organized with row separators $tau_{text{nl}}$ and global demarcations: $$Z_{text{final}} = big[ g_phi(f_v(I_{text{global}})) ; ; tau_{text{split}} ; ; text{GridJoin}big( {g_phi(f_v(I_{i, j}))}, tau_{text{nl}} big) big]$$ Total token budget: $$N_{text{total}} = (m cdot n + 1) times N_{text{crop}} + m cdot N_{text{nl}}$$ For $m=n=2$ and $N_{text{crop}} = 576$, $N_{text{total}} = (4 + 1) times 576 + 2 = 2882$ tokens.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
深度剖析与工程权衡:① ‘全局缩略图不可省’——没有它,模型只见局部(无法理解 tile 间关系);这是 AnyRes 的关键设计。② ’tile 间关系’是 tiling 的核心难题——解法:(a) 全局缩略图、(b) tile 位置嵌入、(c) 让 LLM 的注意力跨 tile 交互(把不同 tile 的 token 拼在一起送 LLM)。③ ‘token 数 ∝ 面积’的代价——3×3 tile + 全局 ≈ 10×N_tile(约 5760 token for 336²),成本高;故需 token 压缩。④ ‘边界切断’的实际影响——物体/文字跨 tile 边界时信息被割裂;解法:(a) 重叠切分、(b) 高分辨率原生(无切分)、(c) 让模型’拼接’(靠 LLM 的跨 tile 注意力)。⑤ ‘复用预训练 ViT’是 tiling 的工程优势——不需改视觉塔(原生分辨率需 ViT 支持可变输入 + 2D RoPE);故 tiling 更易实现(尤其在有现成 CLIP-ViT 时)。⑥ 面试要点——被问’tiling 怎么做’,应给出’切成 k 个固定尺寸 tile + 全局缩略图 + tile 位置编码‘与’token 数 ∝ k(成本高,需压缩)‘,并对比’tiling(复用固定 ViT)vs 原生分辨率(需 2D RoPE)‘;能指出’全局缩略图不可省’是深度理解的标志。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Deep Dive & Engineering Trade-offs: ① The Boundary Cleavage Dilemma: When an object or a line of text falls directly on a tile boundary, half the object resides in Tile $(1,1)$ and half in Tile $(1,2)$. While humans integrate this seamlessly, vision encoders process crops in complete isolation. Without the global overview thumbnail, the LLM struggles to assemble the severed pieces. Adding a small spatial overlap margin (e.g., 10-15% pixel overlap between adjacent tiles) alleviates boundary truncation at the cost of slight token inflation. ② Throughput Bottlenecks in Batch Inference: Slicing an image into 6 tiles increases the effective batch size processed by the vision tower by $7times$. During high-concurrency production serving, large images can monopolize the vision encoder pipeline, causing latency spikes for concurrent lightweight requests. Dynamic tiling managers must enforce hard caps on maximum tiles (e.g., $K_{max} = 6$ or $9$). ③ Dynamic Tile Pruning: In large documents, several tiles often contain empty white margins or solid background colors. Evaluating tile information entropy or background pixel variance allows pruning empty tiles entirely, saving up to 40% of visual tokens on scanned PDFs. ④ Position Encoding Alignment: Each tile must receive 2D positional coordinates reflecting its true global location within the macro image rather than resetting coordinates to $(0,0)$ within each local tile. ⑤ Interview Strategy: Detail the grid calculation formula, explain the local-global token assembly structure with newline tokens, analyze the tile boundary cleavage failure mode, and propose overlap margins and entropy-based tile pruning.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 只切 tile 不加全局缩略图(失去全局视野)
- ⚠️ 忽略 tile 间关系与位置编码
English Pitfalls:
– Resetting 2D positional coordinates to (0,0) inside every local tile, preventing the LLM from understanding inter-tile spatial layout
– Failing to enforce a maximum tile cap ($K_{max}$), allowing high-resolution poster uploads to exhaust GPU memory
– Ignoring tile boundary severance where critical text lines or small objects are split across adjacent independent crops
六、高频深度面试追问与预测 (Follow-Up Questions)
- 为什么需要全局缩略图?
- How does introducing an overlap margin between adjacent image tiles prevent boundary truncation errors in OCR tasks?
- tile 之间的边界如何拼接?
- What criteria can be used to dynamically prune uninformative background tiles before passing them to the language model?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
高分辨率图像切图:LLaVA-NeXT AnyRes 分块、Token 压缩与长图文建模(AnyRes Dynamic Tiling & Visual Token Compression) - 🗺️ 知识图谱模块:
多模态与扩散模型导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。