» Tag
llm-inference
2 posts«August 2026
Apple Silicon macOS VMs Get 11-16x Faster LLM Inference via Llama.cpp
Cua's Metal capability shim delivers 11-16x faster llama.cpp LLM inference in macOS VMs on Apple Silicon, with open benchmarks and source.
Prefill/Decode Disaggregation Can Worsen Tail Latency, Not Fix It
Splitting prefill and decode across GPU pools adds queues and KV transfer overhead that can worsen tail latency without careful control-loop design.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.com