» Tag
machine-learning
185 postsSelf-Play LLM Judges Reward Persuasion, Not Correctness
Self-rewarding LLM judges score plausibility, not correctness; GSM8K experiments show judge approval climbing while true accuracy stays flat, revealing reward hacking.
Why price per 1M tokens is a misleading AI metric
Comparing AI models by price per 1M tokens can mislead teams. Tokenizer differences and chain-of-thought efficiency matter far more than the sticker price per token.
Mnemiq: Open-Source Text-to-SQL Tuned for Your Database
Mnemiq is an open-source text-to-SQL engine that answers natural language queries tailored to your database.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comSparse Policy Selection in RL for LLM Reasoning, Not Capability Learning
Reinforcement learning enhances LLM reasoning by focusing on sparse policy selection rather than teaching new capabilities.
Open Model Enhanced for Scientific Inquiry
The open model developed by Loka and Arcee AI enhances scientific inquiry through tool use and logical reasoning.
NVIDIA Alpamayo 2 Super: 34B Open Vision-Language-Action Model
NVIDIA introduces Alpamayo 2 Super, a 34B open vision-language-action model for robotaxis.
Thinking Machines Launches Inkling Small Open Source AI Model
Thinking Machines introduces Inkling-Small, a compact open-source AI model with impressive performance.
Measuring LLMs’ Ability to Perform Cryptanalysis
A new benchmark evaluates LLMs' cryptanalysis capabilities, revealing new vulnerabilities in cryptographic schemes.
Profiling in PyTorch (Part 3): Attention is All You Profile
Explore profiling the attention mechanism in PyTorch, focusing on performance enhancements through in-place operations.
AI Reverse Engineering Benchmark
AgentRE-Bench assesses AI agents' reverse engineering capabilities. Calibration proves more effective than reasoning depth.