» Tag
vllm
14 postsPorting vLLM's Serving Stack to C++20: A 66 MiB Binary Without Python
The C++20 port of vLLM results in a 66 MiB binary without Python dependencies, achieving comparable speeds at high concurrency.
Llamafile vs vLLM: Two Ways to Serve a Local Model
Explore the differences between Llamafile and vLLM for serving local models and their respective advantages.
Trendyol's agent tunes LLM serving configs for 4x speedup
Trendyol Tech built autooptimizer, an AI agent that autonomously tunes vLLM serving configs, achieving a 4x throughput-latency score gain on Gemma 4 26B with zero manual tweaking.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comDeepseek V4 Flash Achieves ~160 t/s on RTX 6000
Details on achieving ~160 t/s with Deepseek V4 Flash on RTX 6000 and setup information.