核心 GUIDE

PATTERNINTERMEDIATE7 分钟阅读

API Rate-limit Resilience

Rate-limit Resilience 通过 Concurrency、Backoff、Queue、Budget 与 Fallback 管理 Provider 或 Internal Service 的 Quota,而不是让每个 429 都变成无控制 Retry Storm。

核心心智模型

Rate Limit 是 Capacity Signal,不只是 Error。Application 应主动 Shape Demand,让 Work 进入一个 Retry、Priority、Latency 都与可用 Capacity 匹配的 Bounded System。

为什么重要

AI Workload 常具有 Burst 与高 Cost 特征,Parallel Agent、Streaming Call 与 Retry 会比普通 RPS 更快放大流量。没有 Capacity Policy 时,一个小 Quota Event 就可能演变成 Timeout、Duplicate Work 与用户可见不稳定。

01

Retry 前先 Shape Demand

理解 Provider 的关键 Limit 与应用自己的 Concurrency Budget。对 Retryable Response 使用 Bounded Queue、Jittered Exponential Backoff 与显式 Retry Ceiling;按交互价值和 Priority 调度;遵守 Server Retry Hint。只有 Alternative 能保持 Hard Capability Contract 时才 Fallback,否则应该 Queue 或 Explicit Degrade。Telemetry 中记录 Queue Time 与 Throttling。

02

示例:五个 Agent 同时重试同一 Quota Failure

Workflow 在账户刚触达 Per-minute Quota 时启动五个 Parallel Model Call。每个 Worker 立即 Retry,下一时间窗尚未恢复就再次过载。Runtime 改为集中控制 Concurrency、排队低优先任务并使用 Jitter Backoff 后,Capacity 能逐步恢复而不会继续放大请求。

常见失败模式

  • 每个 Worker 遇到 Rate Limit 都立即自行 Retry。
  • 把 Provider Quota 与 Application Concurrency / Queue Policy 分开看。
  • Fallback Model 无法满足同一 Output/Tool Contract,却仍静默切换。

工程启发

  • 集中管理 Concurrency 与 Retry Budget,避免 Worker 自放大。
  • 对 Retryable Throttling 使用 Jitter、Retry Ceiling 与 Server Hint。
  • 把 Queue Time、Throttled Request、Fallback Outcome 当成 Reliability Signal。

关键结论

  1. 01Rate Limit 应改变 Demand Shape,而不是触发 Blind Retry。
  2. 02Concurrency Control 是 LLM Application Architecture 的一部分。
  3. 03Fallback 必须保留 Hard Contract,否则显式降级。

阅读记录

未打开Practice 尚未完成

这里只记录你真实做过的动作,不代表掌握、熟练或认证。

用于这些学习路径

这个 Concept 会在多个 canonical 学习路径中复用。