Discrete-event simulation for LLM serving capacity planning: how many GPUs for a p99 TTFT SLO under ShareGPT traffic, batching fragmentation, and autoscaling lag — with cliff analysis and three levels of planning rigor.
performance-engineering benchmark simulation latency capacity-planning inference systems slo autoscaling cost-optimization discrete-event-simulation kv-cache llm sharegpt llm-serving batch-scheduling machine-learning-infrastructure serving-systems gpu-capacity
-
Updated
Jul 26, 2026 - Python