« All posts

Bursty Arrivals Accelerate LLM Inference Times

Bursty workloads have been found to unexpectedly speed up LLM inference.

Bursty workloads have led to unexpected speed increases in LLM inference. The burstiness allows for the separation of tokens from large prefills, reducing interference. However, this effect is largely attributed to a kernel optimization artifact rather than a fundamental change in workload behavior.