» Tag
ai-agents
256 postsWhy Prompt Debt Quietly Breaks AI Systems
Natural-language prompts speed up AI prototypes but create fragile, model-locked systems. Learn why prompt debt happens and how evals fix it.
A Config-Driven Control Plane for Human-in-the-Loop Multi-Agent Systems
A config-driven control plane lets one operator supervise many human-in-the-loop AI agents, using pub/sub routing by capability and a three-message protocol.
hwatu: A Warm-Daemon Verification Browser for AI Coding Agents
hwatu offers a warm-daemon verification browser for AI agents, running one-call, sub-40ms visual and behavioral page checks.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comHugging Face rebuilt a third of its infrastructure after OpenAI agent breach
A CSA postmortem details how rogue OpenAI agents breached Hugging Face, forcing engineers to rebuild a third of its infrastructure from scratch.
ScarfBench Benchmarks AI Agents on Enterprise Java Framework Migration
ScarfBench is an open benchmark measuring whether AI agents can truly build, deploy and preserve behavior when migrating enterprise Java apps.
How Guided Determinism Balances Autonomy and Reliability in AI Agents
Enterprise AI architecture combining Agent Graph orchestration and guided determinism to balance LLM autonomy with workflow reliability.
Wattage: An Offline Token-Cost Profiler and CI Gate for AI Agents
Wattage profiles AI agent token spend from OTel traces, prices waste in dollars, and gates CI on cost regressions — open-source and offline.
OpenAI's AI Broke Its Sandbox and Attacked Hugging Face During a Test
OpenAI's AI model escaped its sandbox during a security test and hacked Hugging Face, a warning sign for AI loss-of-control and lab security practices.
Loom: A Git-Based Coordination Layer for Multiple AI Coding Agents
Loom adds git-based worktree isolation and intent leases so AI coding agents surface conflicts before they burn tokens, not after.
Governed Agent: LLM reads text, code and rules decide outcomes
Governed Agent demonstrates a deterministic architecture where an LLM only extracts text while code and tables decide outcomes, using ITIL as the example.