» Tag
llm
535 postsGenRec: An LLM-Backed Recommendation Ranker at Netflix
Netflix's GenRec leverages LLMs to enhance recommendation systems, achieving significant performance improvements with fewer training examples.
Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation
This study explores how prior scores in LLM-as-a-Judge systems affect evaluations, emphasizing the need for careful context engineering.
Reconstructing the Benchmark Behind Luc Julia's 64% LLM Reliability Claim
A new resource reconstructs the basis of Luc Julia's 64% reliability claim for LLMs.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comEnhance DeepSeek Harness with LLM-as-a-Verifier Plugin
The LLM verifier plugin for DeepSeek Harness grades candidate solutions and returns the best one.
Programmatic Memory for Long-Horizon LLM Agents
PRO-LONG enhances long-horizon LLM agents' performance by integrating programmatic memory.
Agentic Engineering Applications at Zalando
Zalando enhances Agentic Engineering with LLMs. LiteLLM-based proxy allows engineers to experiment with various tools.
Vale-LLM-slop: Prose Linting for LLMs
Vale-LLM-slop aids in prose linting for LLMs, ensuring clearer and more understandable content.
Why Prompt Injection Remains Possible in LLM Applications
Prompt injection is still a valid threat in LLM applications. Discover why in this article.
Testing Real-Time Verification of LLM Cache Hits: A Weak Yes
Results of an empirical study on real-time verification of LLM cache hits.
Integrating Scikit-Ollama with Scikit-LLM/Ollama
Scikit-ollama integrates scikit-learn with local Ollama models for zero-shot text classification.