From LLM Inference to Agentic Workloads: Characterization and Implications
Explore how agentic applications are reshaping LLM services and their system behaviors.
Agentic applications are transforming AI services from isolated model inference to long-running workloads where LLMs coordinate tools, environments, and persistent state. However, the system behavior of these workloads—especially regarding latency, cost, and bottlenecks—remains poorly characterized. The introduction of AgentSysBench provides a benchmark suite with ten representative agentic applications, revealing actionable insights that can significantly enhance performance, such as reducing latency by up to 40%.
This synthesis was produced from its source by AI; there is no human editor or manual review step. How we work