» Tag
evaluation
42 postsEvaluating Open-Source AI Tools: A Structured Approach
Adopt a structured approach for evaluating open-source AI tools, focusing on workflow fit, security, and reproducibility of results.
LLM-Generated Tests Influence Implementation Selection Outcomes
Key findings on the impact of LLM-generated tests on implementation selection and evaluation.
Comparing LLM Evaluation Frameworks: Measuring Model Performance
Learn to evaluate LLM applications with RAGAS, DeepEval, and Promptfoo. Address biases effectively in your designs.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comPoPE: Evaluating Error-Conditioned Self-Repair in Small Code LLMs
PoPE introduces a new standard for evaluating self-repair in small code LLMs, emphasizing rigorous placebo-controlled assessments.
How to Evaluate a Slide-Change Detector Effectively
Steps and key considerations for effectively evaluating slide-change detectors.
Agentic Tool-Use Evaluation on Local 35B Model: Insights and Challenges
Insights from an agentic evaluation on a local 35B model, focusing on tool use and performance.
Evaluating AI Engineers in 2026: A 7-Point Framework
A 7-point framework for evaluating AI engineers in 2026. Identifying the right role and key factors is essential.
Evaluating AI Agent Skill Performance with NVIDIA SkillEvaluator
NVIDIA SkillEvaluator is an open-source tool for measuring AI agent skill performance.
Eval-Driven Development: Lessons from Evaluating GenAI at Scale
Airbnb prioritizes evaluation in Generative AI product development, sharing best practices for engineers.
Building a Practical Taxonomy for AI World Models
A report addressing the definition and classification of world models. It offers important insights for engineers.