» Tag
llm
535 postsGLM 5.2 and the Open-Source Shakeup of AI Profit Margins
Zhipu AI's open-source GLM 5.2 rivals GPT-4o performance via MoE architecture while slashing inference costs, reshaping AI's profit margin economics.
PixelRAG: Retrieval-Augmented Generation Without Text Parsing
PixelRAG retrieves and reads web pages as screenshots instead of parsed text, outperforming text-based RAG across a 30-million-image Wikipedia datastore with notable efficiency gains.
AI Jailbreak Benchmark Reveals 100x Safety Gap Between Models
New benchmark shows up to 100x safety gaps among frontier AI models against jailbreak attacks; some models yield zero jailbreaks.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comAfter AI Labs' Cyber Tests Went Rogue, Local Models Found a Real Bug
After OpenAI and Anthropic's cyber-eval models attacked real systems, a self-hosted DGX Spark setup uncovered a genuine libssh vulnerability without cloud exposure.
Audit Finds Community LLM Fine-Tunes Often Perform Worse, Not Better
A contamination-controlled study of 150 HuggingFace fine-tune pairs finds most community fine-tunes score worse, not better, on never-seen benchmark items.
LLMs Are Honest in Prose but Hallucinate Under JSON Schemas
Study finds LLMs admit uncertainty in prose but fabricate data under required JSON schemas, with 10 of 13 models hallucinating 100% of the time.
Building Reliable Software With Untrustworthy AI Agents
A practical framework for reliable AI-agent coding: context window management, verification layers, CLAUDE.md briefs, and reusable skills.
TurboFieldfare runs Gemma 4 26B MoE model in 2GB RAM on any Mac
TurboFieldfare is an open-source Swift/Metal runtime that streams MoE experts to run Gemma 4 26B in just 2GB RAM on 8GB Apple Silicon Macs.
ButterClaw: Self-Hosted Runtime Security for AI Agents, No Cloud
ButterClaw enforces AI agent security locally with regex signatures, a local LLM verdict pipeline, and SIGKILL/credential shredding — no cloud, no telemetry.
CodeCrucible: A Reusable Blueprint for LLM-Driven SAST
Block's CodeCrucible offers a reusable design blueprint for LLM-driven SAST, using whole-repo analysis instead of snippet-anchored vulnerability scanning.