» Tag
mechanistic-interpretability
4 postsNo Fine-Tuning: Facts Hand-Wired Directly Into Llama-3.1-8B's Weights
A mechanistic-interpretability method hand-wires facts into Llama-3.1-8B's weights without fine-tuning, LoRA, or RAG — with a live neuron visualizer.
ModelMRI: A Local Debugger for Peering Inside LLMs, VLMs and Robot Policies
ModelMRI is a local, open-source tool for inspecting attention, activation patching and concepts inside LLMs, VLMs and robot policies.
Mechanistic View Reveals How Bias Lives Inside LLM Judges
Study shows LLM-as-judge bias is encoded in activation geometry, enabling causal steering and better failure prediction than text-based methods.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comNew method flags risky tool calls in AI agents before they happen
Researchers built a sparse-autoencoder and probe-based toolkit that reads AI agent internals to flag risky or unnecessary tool calls before execution happens.