» Tag
inference
44 postsDarkbloom: AI Inference on Macs and Security Audit Results
Darkbloom enables AI inference on Macs. Discover the results of the recent security audit and findings.
Unified Memory: Why Mini PCs Run 70B Models a Big GPU Can't
Unified-memory mini PCs like AMD's Strix Halo can load 70B-parameter models that a $2,000 RTX 5090 cannot fit. Here's why capacity and bandwidth pull in opposite directions.
OpenAI and Broadcom Unveil Jalapeño, an LLM Inference Chip
OpenAI and Broadcom unveiled Jalapeño, a custom LLM inference chip built in nine months, set for gigawatt-scale data center deployment starting in 2026.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comNvidia and Cerebras: Selling Performance Customers Will Likely Never See
Nvidia and Cerebras showcased their performance metrics at Hot Chips. However, these figures may not reflect real-world usage.
WISP: A CUDA Engine for Streaming 744B+ Parameter MoE Models on Consumer Hardware
WISP is a CUDA engine for streaming 744B+ parameter MoE models on consumer hardware.
Ninfer: High-performance single-GPU inference
NInfer offers a high-performance C++/CUDA inference engine for RTX 5090.
The State of Open-Source LLM Inference
Open-source LLM inference is a key topic for businesses regarding cost and data control. This entry discusses strategies and challenges.
Inside vLLM: Anatomy of a High-Throughput LLM Inference System
Explore the core components and features of vLLM's high-throughput LLM inference system.
Co-Designing AI Model Attention for Fast, Interactive Long-Context Inference
Explore practical guidelines for optimizing AI model attention and inference efficiency.
Final Token Preference Optimization Tackles Reasoning Model Doom Loops
Antidoom uses Final Token Preference Optimization to fix repetitive doom loops in reasoning models, cutting loop rates sharply in LFM2.5 and Qwen3.5 without broad model degradation.