Inference
2 pages
-
AI Engineering
source
*AI Engineering* — Chip Huyen
-
Inference Optimization
concept
Inference metrics (TTFT, TPOT, throughput, goodput, MFU, MBU); prefill (compute-bound) vs decode (memory bandwidth-bound); speculative decoding; KV cache management (PagedAttention, FlashAttention, GQA/MQA); continuous batching; prefill-decode decoupling; prompt caching