» Tag
evaluation
46 postsYour RAG Eval Is Checking the Receipt, Not the Patient
Explore the risks of entity attribution failures in clinical RAG systems and how to improve evaluations.
Evaluating Open-Source AI Tools: A Structured Approach
Adopt a structured approach for evaluating open-source AI tools, focusing on workflow fit, security, and reproducibility of results.
Replacing LLM Judges with Jev Decision Models
A system using Jev decision models replaces LLM judges, offering fast evals at low cost.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comLLM-Generated Tests Influence Implementation Selection Outcomes
Key findings on the impact of LLM-generated tests on implementation selection and evaluation.
Comparing LLM Evaluation Frameworks: Measuring Model Performance
Learn to evaluate LLM applications with RAGAS, DeepEval, and Promptfoo. Address biases effectively in your designs.
PoPE: Evaluating Error-Conditioned Self-Repair in Small Code LLMs
PoPE introduces a new standard for evaluating self-repair in small code LLMs, emphasizing rigorous placebo-controlled assessments.
How to Evaluate a Slide-Change Detector Effectively
Steps and key considerations for effectively evaluating slide-change detectors.
Agentic Tool-Use Evaluation on Local 35B Model: Insights and Challenges
Insights from an agentic evaluation on a local 35B model, focusing on tool use and performance.
Evaluating AI Engineers in 2026: A 7-Point Framework
A 7-point framework for evaluating AI engineers in 2026. Identifying the right role and key factors is essential.
Evaluating AI Agent Skill Performance with NVIDIA SkillEvaluator
NVIDIA SkillEvaluator is an open-source tool for measuring AI agent skill performance.