» Tag
inference
44 postsPorting vLLM's Serving Stack to C++20: A 66 MiB Binary Without Python
The C++20 port of vLLM results in a 66 MiB binary without Python dependencies, achieving comparable speeds at high concurrency.
Full-fabric VHDL LLM Inference Engine: Runs Qwen3.5-class
The full-fabric VHDL LLM inference engine runs Qwen3.5-class transformer inference on FPGA.
AMD Acquires Taalas to Enhance AI Chip Inference Performance
AMD's acquisition of Taalas aims to boost AI chip performance and compete with Nvidia.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comBursty Arrivals Accelerate LLM Inference Times
Bursty workloads have been found to unexpectedly speed up LLM inference.
At-the-Roofline Sparse Tensor Contractions on Vector Processors for Inference
Ventaglio enhances sparse tensor contractions on vector processors, boosting inference performance significantly.
Routing LLM Traffic with TCP-Style Congestion Control
Explore an adaptive router that optimizes LLM traffic based on cost and speed.
Assessing LLM Inference Profitability: Insights from Kimi K3
Analysis and calculations on the profitability of LLM inference using Kimi K3.
We Built a Day-0 API for Kimi K3
Baseten has introduced day-0 support for Kimi K3 on its Model APIs, addressing the challenges of its 2.8T parameters.
Introduction to LLM Inference: Process and Its Importance
Explore the LLM inference process and its significance in engineering. Gain insights into performance and model structure.
RAI: CPU-only LLM Inference Engine Built in Pure Rust
RAI is a CPU-only LLM inference engine in Rust, providing high performance without GPU dependencies.