Skip to Content
人工智能系统性能工程 (Chinese Edition)
book

人工智能系统性能工程 (Chinese Edition)

by Chris Fregly
November 2025
Intermediate to advanced
1060 pages
14h 20m
Chinese
O'Reilly Media, Inc.
Content preview from 人工智能系统性能工程 (Chinese Edition)

第15章 多节点 推理、并行性、解码与路由优化

本作品已使用人工智能进行翻译。欢迎您提供反馈和意见:translation-feedback@oreilly.com

大型语言模型(LLMs)的参数规模持续攀升至庞大的 量级。尤其随着混合专家模型(MoE)的出现——这类模型通过内置专家门控机制整合众多专业子网络("专家")——模型参数规模已突破数百亿乃至数万亿量级。尽管实际处理特定输入时仅需调用其中一小部分参数,但对如此庞大的模型进行推理仍需将工作负载分布到多块GPU上。

本章聚焦于利用现代NVIDIA GPU为这些庞大LLMs实现高效高性能多节点推理的高级优化技术。我们将探讨如何通过架构师分布式推理系统以最小化延迟并最大化吞吐量——同时运用硬件与算法创新。

首先探讨解耦预填充与解码(PD,解耦PD)架构 ,该架构将推理工作负载拆分为可独立调优的独立阶段。随后深入解析数据、张量、流水线、专家模型及上下文等核心推理并行策略,阐述如何组合运用这些策略在多GPU环境中服务大型模型。

随后我们将介绍投机性解码方法,包括Medusa、EAGLE及草稿验证方案等技术。这些方法允许在推理过程中生成并评估多个令牌,突破传统自回归LLMs单令牌生成的限制,从而克服顺序解码瓶颈。我们还将探讨约束解码技术(如定制JSON模式)以强制输出格式规范,以及面向MoE模型的动态路由策略,从而提升系统的专家门控与负载均衡效率。

解耦预填充与解码架构

如前所述,现代LLMs的推理流程 包含预填充与解码两个阶段。通过实现解耦式预填充与解码,可将这两个阶段分离。这使得预填充集群与解码集群能够独立扩展——甚至可在不同硬件平台上运行——从而显著提升大规模LLM服务的性能,本章后文将详细阐述。

跨供应商或跨架构部署要求两端KV缓存布局和数据类型保持一致。实际生产系统应将预填充和解码操作部署在兼容的GPU家族上,通过统一数值格式实现KV缓存转移与数据复用。

在预填充阶段,模型通过单次前向传播处理整个输入prompt(通常包含数千、数万甚至数百万个词元),生成由LLM计算的初始隐藏状态。随后为输入prompt中的所有词元填充注意力键值(KV)缓存。图15-1展示了分离式预填充与解码如何共享KV缓存,并实现KV传输与计算的重叠执行。

Diagram illustrating the overlap of KV cache transfers with computations during prefill and decode stages, showing that transfer affects tokens 1 and 2 only.
图15-1. 分解式 预填充与解码共享KV缓存,实现KV传输与计算的重叠

解码阶段中,模型通过自回归生成机制预测序列中的每个新令牌。该过程需调用所有先前生成令牌的缓存注意力KV表示。

投机性解码通过在单批次中预生成多个令牌来加速 解码过程。随后并行验证令牌正确性,从而降低标准逐令牌自回归解码的顺序性。投机性解码将在后续章节详述。

预填充-解码干扰

传统上,LLM推理系统将 这两个阶段部署在相同节点上,并简单地将所有计算批量处理。然而,这种简单粗暴的方法导致了所谓的预填充-解码干扰。例如,长时间的prompt预填充会占用GPU资源,延迟其他请求的时敏解码工作——反之亦然。

将预填充与解码置于同一节点,迫使系统采用单一的调度和资源分配策略处理这两种特性截然不同的阶段:预填充涉及大规模并行计算,而解码则需要大量小规模顺序计算。结果导致系统必须在优先保障某一阶段性能与过度配置硬件以满足双重需求之间做出取舍。

采用解耦预填充与解码架构时, 预填充和解码阶段被分配至不同GPU池,从而消除两者工作负载的直接干扰。 ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

AirBnbBlueOriginElectronic ArtsHomeDepotNasdaqRakutenTata Consultancy Services

QuotationMarkO’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.
Julian F.
Head of Cybersecurity
QuotationMarkI wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.
Addison B.
Field Engineer
QuotationMarkI’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.
Amir M.
Data Platform Tech Lead
QuotationMarkI'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.
Mark W.
Embedded Software Engineer

You might also like

向量数据库 (Chinese Edition)

向量数据库 (Chinese Edition)

Nitin Borwankar

Publisher Resources

ISBN: 0642572281557