» Tag
vllm
3 posts«August 2026
Single MI300X, Real Coding Agents: DeepSeek V4 Flash Throughput Tested
Real-world benchmark of DeepSeek V4 Flash on a single MI300X GPU serving coding agents, covering throughput, caching, and cost per token.
Inside vLLM: Anatomy of a High-Throughput LLM Inference System
Explore the core components and features of vLLM's high-throughput LLM inference system.
Porting vLLM's Serving Stack to C++20: A 66 MiB Binary Without Python
The C++20 port of vLLM results in a 66 MiB binary without Python dependencies, achieving comparable speeds at high concurrency.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.com