核心 GUIDE

METRICFOUNDATION7 分钟阅读

Latency 与 Throughput 基础

AI Serving Performance 本质上是 Queueing 与 Workload 问题:响应时间、并发和 Token Generation 会共同竞争有限 Compute Capacity。

核心心智模型

Latency 描述单个 Request 等待并执行多久;Throughput 描述系统单位时间完成多少 Work。Batching、Concurrency 与 Model Size 共享 Capacity,因此优化一个指标可能伤害另一个。

为什么重要

一个 Model 在单请求 Demo 里很快,放到并发流量下可能完全不可用。Long Prompt 增加 Prefill Work,Long Generation 长时间占据 Serving Capacity,而 Aggressive Batching 虽然提升 Throughput,却可能让单个请求等待更久。Product Architecture 应明确 TTFT、Generation Speed、Tail Latency 与 Concurrency Target,而不是只写一句“足够快”。

01

把 Response Time 拆成 Workload Component

分别测 Queue Wait、Input Processing / Prefill、First-token Delay 与后续 Token Generation。在 Representative Prompt Length 和 Concurrent Load 下测量,而不是只测单请求。之后再根据 Product Objective 选择 Smaller Model、Shorter Context、Caching、Batching、Routing 或 Admission Control 等 Trade-off。

02

例子:流量高峰下的 Support Chat

低负载时,大模型不到一秒就开始输出;流量高峰时,请求进入 Queue,Tail Latency 超过十秒,即使每 Token Generation Speed 并没有变。真正修复可能是限制 Concurrency、把 Routine Case Route 到 Smaller Model,或减少 Context Cost,而不是改 Prompt 文案。

常见失败模式

  • 只用单请求 Benchmark Latency,再直接外推 Production Concurrency。
  • 只优化 Average Latency,忽略用户真正感知到的 Tail Behavior。
  • 增加 Context 或 Output Length,却不计算它对 Shared Serving Capacity 的影响。

工程启发

  • Streaming Product 分别记录 Time-to-first-token 与 End-to-end Latency。
  • 用 Representative Prompt Length 与 Concurrency Level 做压力测试。
  • 把 Cost 与 Successful-task Throughput 与 Tokens-per-second 一起看。

关键结论

  1. 01Serving Speed 取决于 Workload 与 Queueing,不只取决于 Model Architecture。
  2. 02Throughput 优化有时会提高 Individual Latency。
  3. 03Capacity Target 应表达成 Product-level Service Objective。

用于这些学习路径

这个 Concept 会在多个 canonical 学习路径中复用。