» Tag
evaluation
46 postsEval-Driven Development: Lessons from Evaluating GenAI at Scale
Airbnb prioritizes evaluation in Generative AI product development, sharing best practices for engineers.
Building a Practical Taxonomy for AI World Models
A report addressing the definition and classification of world models. It offers important insights for engineers.
Automating Evaluation Design and Hillclimbing with Claude
The Claude API skill offers tools for designing evaluations and improving application performance.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comAnchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation
This study explores how prior scores in LLM-as-a-Judge systems affect evaluations, emphasizing the need for careful context engineering.
My Local LLM Scored 6/6 but Was Wrong Every Time
Discover the difference between answer format and value in local LLM evaluations.
SELECT first_name, last_name vs SELECT first_name || ' ' || last_name: Accuracy Issues
Explore the impact of column concatenation on SQL evaluation accuracy and solutions.