» Tag
machine-learning
25 postsCo-Designing AI Model Attention for Fast, Interactive Long-Context Inference
Explore practical guidelines for optimizing AI model attention and inference efficiency.
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
AutoDesign optimizes long-horizon design with a meta-harness, enhancing human preferences in media output.
Filtered Vector Search: What Acorn Fixes, and What Fixes Acorn
Qdrant introduces new methods to solve filtered vector search issues. ACORN and filterable HNSW enhance performance.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comNo cloud, no GPUs: Liquid AI's LFM2.5-2.6B model empowers edge devices
Liquid AI introduces LFM2.5-2.6B, a model that runs on local hardware without cloud or GPU dependencies.
NVIDIA VoiceChat-11B: AI Speech Model for Real-Time Communication
NVIDIA VoiceChat-11B is an AI speech model with 11 billion parameters, designed for real-time voice interactions.
Solving Moe Load Imbalance in LLM Training via Optimal Transport
TAOT method improves MoE training speed by 43% while reducing communication costs by 74%.
1.5B Model Trained to Write Shell Commands on a Laptop CPU
User trains a 1.5B Qwen2.5-Coder model to generate shell commands. It operates on a laptop.
Hollow-LLM Attack: Ghost Weights That Fool Zero-Knowledge LLM Verification
The Hollow-LLM attack examines the impact of ghost weights on zero-knowledge verification.
Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
The v6.0 update of Sentence Transformers introduces MultiVectorEncoder for enhanced late interaction retrieval.
Benchmarking Cheap LLMs for Production Agent Traces
We benchmarked cheaper LLMs for production agent traces. The results were significant.