» Tag
llama.cpp
13 postsNilaMind: A Fully On-Device 1.5B LLM for Mental Health Support
NilaMind runs Qwen2.5-1.5B via llama.cpp fully offline on Android, with crisis safety enforced by a deterministic, model-independent gate.
Support for GBNF Grammar Added for Local Models
Support for GBNF grammar has been added for local models, ensuring strict output validation.
DFlash in llama.cpp: 4.44x Faster Local Inference on Qwen 3.6 27B
DFlash, merged into llama.cpp, uses block-diffusion drafting to boost Qwen 3.6 27B inference speed, hitting 4.44x at 36K context with near-lossless quality.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.com