» Tag
llm
535 postsNVIDIA's Nemotron 3 Super Beats GPT-OSS-120B on Coding Benchmarks
NVIDIA's open Nemotron 3 Super model beats GPT-OSS-120B by 20 points on SWE-Bench and delivers 2.2x faster inference throughput.
Open-source AI skill set turns LinkedIn posting into a pipeline
linkedin-skills is an open-source, ten-skill toolkit for Claude Code and Codex that automates LinkedIn writing, publishing, and analytics safely.
10 AI Coding Models, 5 Tasks: Price Doesn't Predict Quality
Benchmarking 10 LLMs across 5 coding tasks reveals price and code quality barely correlate, with budget models rivaling premium ones.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.com"Hallucination" Isn't One Bug. It's Three, and Only One Is Fixable
Hallucination isn't one failure mode — it's three. A test to tell them apart, why benchmarks reward bluffing, and what a 2026 prediction experiment showed.
Why $/Token Pricing Hides the Real Cost of Frontier AI Models
Frontier AI pricing pages hide tokenizer differences that can inflate real costs by up to 73% on code like TypeScript, per new billing analysis.
AEGIS: An Open-Source, Self-Hosted Personal AI Orchestration System
Developer open-sources AEGIS, an MIT-licensed, self-hosted personal AI orchestration platform built on FastAPI, Postgres, and Temporal.
VetoBench Tests Whether AI Agent Memory Retains Rejected Decisions
VetoBench is an open benchmark testing whether AI agent memory systems retain and surface previously rejected technical decisions.
Claude Code Burns 33K Tokens Before It Even Reads Your Prompt
Wire-level analysis shows Claude Code sends 33K tokens before reading your prompt - 4.7x OpenCode, with 3.7x higher real-task costs.
Hermes Agent: A Self-Improving AI Framework With Persistent Memory
Nous Research's open-source Hermes Agent framework combines persistent memory, reusable skills, and a real multi-agent architecture for self-improving AI.
Benchmarking a Markdown Knowledge Graph as AI Agent Memory
IWE tested markdown knowledge graphs as AI agent memory using the LOCOMO benchmark, reaching 96% of a hand-built ceiling with a cheap curator model.