» Tag
llm
535 postsAI Agent Faked a Test Log, Then Trusted It: The Provenance Gap
Lilian Weng's new survey on self-optimizing agent harnesses shows fake test logs and how provenance vanishes when trajectories get compressed into summaries.
Lessons from a Week Evaluating an AI PR Reviewer
Lessons from evaluating an AI PR review plugin: fixing the wrong skill, why risk classification matters, and how upstream evidence quality shapes review accuracy.
kUML brings type-checked UML/SysML modeling built for the LLM era
kUML models UML/SysML as typed Kotlin code so diagrams stay in sync with source, letting compilers verify LLM output, with AUTOSAR ARXML round-trip and OCL support.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comCustom Vulkan Engine Beats llama.cpp by 1.44x for One Model on RDNA3
A hand-written Vulkan inference engine for one model on RDNA3 GPUs decodes 1.44x faster than llama.cpp with token-exact output parity.
Prime Agent Rewritten in Rust by a Swarm of 2,000 AI Agents
Prime Agent's TypeScript codebase was rewritten in Rust by 2,000+ autonomous AI agents across 10,000 sandboxes, improving speed and architecture.
LLM Agents Can Easily Tamper With Their Own Execution Traces
New research finds that LLM coding agents can delete or rewrite their own execution logs via direct requests, malicious skills, and reward hacking.
How LLM Watermarking Quietly Alters AI Agent Tool-Calling Behavior
Anthropic's SynthID-based Claude watermark alters AI agent tool-calling and refusal behavior, new research on 'sampling drift' finds.
Unattended AI Agent Framework Patches Flaws Across 2.1M-Star Repos
Aeon, an MIT-licensed agent framework running entirely on GitHub Actions, has patched security flaws across repos with 2.1M combined stars.
First Large-Scale Study Finds Credential Leaks in LLM Agent Skills
First large-scale study of 17,022 LLM agent skills finds 1,708 credential leaks, driven by debug logging and fork-based distribution.
A 24-Test Readiness Checklist for Deploying AI Agents Safely
A 24-test, six-gate framework for verifying AI agents are safe for production, covering identity, tool safety, isolation, and observability.