» Tag
inference
33 postsThe State of Open-Source LLM Inference
Open-source LLM inference is a key topic for businesses regarding cost and data control. This entry discusses strategies and challenges.
Inside vLLM: Anatomy of a High-Throughput LLM Inference System
Explore the core components and features of vLLM's high-throughput LLM inference system.
Co-Designing AI Model Attention for Fast, Interactive Long-Context Inference
Explore practical guidelines for optimizing AI model attention and inference efficiency.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comFinal Token Preference Optimization Tackles Reasoning Model Doom Loops
Antidoom uses Final Token Preference Optimization to fix repetitive doom loops in reasoning models, cutting loop rates sharply in LFM2.5 and Qwen3.5 without broad model degradation.
Porting vLLM's Serving Stack to C++20: A 66 MiB Binary Without Python
The C++20 port of vLLM results in a 66 MiB binary without Python dependencies, achieving comparable speeds at high concurrency.
AMD Acquires Taalas to Enhance AI Chip Inference Performance
AMD's acquisition of Taalas aims to boost AI chip performance and compete with Nvidia.
Bursty Arrivals Accelerate LLM Inference Times
Bursty workloads have been found to unexpectedly speed up LLM inference.
At-the-Roofline Sparse Tensor Contractions on Vector Processors for Inference
Ventaglio enhances sparse tensor contractions on vector processors, boosting inference performance significantly.
Assessing LLM Inference Profitability: Insights from Kimi K3
Analysis and calculations on the profitability of LLM inference using Kimi K3.
We Built a Day-0 API for Kimi K3
Baseten has introduced day-0 support for Kimi K3 on its Model APIs, addressing the challenges of its 2.8T parameters.