核心 GUIDE
Latency 与 Throughput 基础
AI Serving Performance 本质上是 Queueing 与 Workload 问题:响应时间、并发和 Token Generation 会共同竞争有限 Compute Capacity。
核心心智模型
Latency 描述单个 Request 等待并执行多久;Throughput 描述系统单位时间完成多少 Work。Batching、Concurrency 与 Model Size 共享 Capacity,因此优化一个指标可能伤害另一个。
为什么重要
一个 Model 在单请求 Demo 里很快,放到并发流量下可能完全不可用。Long Prompt 增加 Prefill Work,Long Generation 长时间占据 Serving Capacity,而 Aggressive Batching 虽然提升 Throughput,却可能让单个请求等待更久。Product Architecture 应明确 TTFT、Generation Speed、Tail Latency 与 Concurrency Target,而不是只写一句“足够快”。
01
把 Response Time 拆成 Workload Component
分别测 Queue Wait、Input Processing / Prefill、First-token Delay 与后续 Token Generation。在 Representative Prompt Length 和 Concurrent Load 下测量,而不是只测单请求。之后再根据 Product Objective 选择 Smaller Model、Shorter Context、Caching、Batching、Routing 或 Admission Control 等 Trade-off。
02
例子:流量高峰下的 Support Chat
低负载时,大模型不到一秒就开始输出;流量高峰时,请求进入 Queue,Tail Latency 超过十秒,即使每 Token Generation Speed 并没有变。真正修复可能是限制 Concurrency、把 Routine Case Route 到 Smaller Model,或减少 Context Cost,而不是改 Prompt 文案。
常见失败模式
- 只用单请求 Benchmark Latency,再直接外推 Production Concurrency。
- 只优化 Average Latency,忽略用户真正感知到的 Tail Behavior。
- 增加 Context 或 Output Length,却不计算它对 Shared Serving Capacity 的影响。
工程启发
- Streaming Product 分别记录 Time-to-first-token 与 End-to-end Latency。
- 用 Representative Prompt Length 与 Concurrency Level 做压力测试。
- 把 Cost 与 Successful-task Throughput 与 Tokens-per-second 一起看。
关键结论
- 01Serving Speed 取决于 Workload 与 Queueing,不只取决于 Model Architecture。
- 02Throughput 优化有时会提高 Individual Latency。
- 03Capacity Target 应表达成 Product-level Service Objective。
用于这些学习路径
这个 Concept 会在多个 canonical 学习路径中复用。