« All posts

Benchmarking Cheap LLMs for Production Agent Traces

We benchmarked cheaper LLMs for production agent traces. The results were significant.

We perform one LLM call on every agent trace, converting them into concise, searchable digests. This call is the fastest-growing cost in our usage. We tested whether a cheaper model could achieve the same results and found a model that could replace Claude Sonnet 4.6 at a significantly lower cost.