Discrete-event simulation benchmark for decode preemption policies in LLM serving: when to interrupt a running decode to admit an urgent prefill, and whether checkpoint or recompute is cheaper.
-
Updated
Jul 26, 2026 - Python
Discrete-event simulation benchmark for decode preemption policies in LLM serving: when to interrupt a running decode to admit an urgent prefill, and whether checkpoint or recompute is cheaper.
Add a description, image, and links to the recompute topic page so that developers can more easily learn about it.
To associate your repository with the recompute topic, visit your repo's landing page and select "manage topics."