» Tag
local-inference
4 postsRunning Kimi K3, a 2.8T-Parameter MoE Model, on an M1 Mac
Deltafin runs Kimi K3, a 2.8T-parameter MoE model, on a 64GB M1 Mac via full local install or expert streaming, no cluster required.
DeepSeek V4 Flash reaches 32 tok/s on AMD Ryzen AI MAX+ 395
AMD Ryzen AI MAX+ 395 runs 284B-parameter DeepSeek V4 Flash locally at 32 tok/s decode and ~250 tok/s sparse prefill using 128GB unified memory.
Qwisp: MoE expert-streaming engine runs Qwen3.6-35B-A3B on 8GB Macs
Qwisp streams MoE experts from flash to run the 35B-parameter Qwen3.6-A3B model on 8GB Macs, with bit-exact lossless decoding and raw-Metal speed.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comCamelid: Local AI Inference in Rust with Multiple Interfaces
Camelid enables local AI inference in Rust with multiple interfaces, running GGUF models directly.