» Tag
llm
478 postsWhy price per 1M tokens is a misleading AI metric
Comparing AI models by price per 1M tokens can mislead teams. Tokenizer differences and chain-of-thought efficiency matter far more than the sticker price per token.
Restore LLM Inference Capacity Quickly with Shadow Engine Recovery in NVIDIA Dynamo
NVIDIA Dynamo's shadow engine recovery allows rapid restoration of LLM processes after failures.
Cross-vendor byte-identical inference for a 72B LLM (AMD MI300X vs. Nvidia H100)
A new protocol enables byte-identical outputs from a 72B LLM using AMD MI300X and Nvidia H100.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comStealing Reasoning Traces from Proprietary LLM APIs
We found a way to extract reasoning traces from frontier AI APIs, highlighting significant security risks and potential data leaks for engineers.
Microsoft's Three-Layer LLM Routing Architecture for AI Agents
Microsoft has unveiled a three-layer LLM routing architecture for Azure Kubernetes Service, designed to efficiently manage agent traffic.
Persistent State Machine: Breaking the von Neumann Memory Wall for LLM Attention
The Persistent State Machine addresses the memory wall in LLMs, cutting energy use and data traffic significantly.
Making LLM Extraction Trustworthy Enough to Act On
Honeycomb addresses the verification processes needed to enhance LLM extraction trustworthiness.
Statgate: Statistically Calibrated Gates for LLM Evaluations
Statgate offers statistically calibrated gates that reduce false positives in LLM evaluations.
Open-ultra: A Self-Training LLM Routing Proxy
Open-ultra is a self-training LLM routing proxy that reduces costs while maintaining high output quality.
Add a Free LLM to a Static Site with Cloudflare Workers AI
Learn how to integrate AI into a static site using Cloudflare Workers AI without a backend or API key, and get a PRD outline in 15 seconds.