» Tag
gpu
82 postsDFlash in llama.cpp: 4.44x Faster Local Inference on Qwen 3.6 27B
DFlash, merged into llama.cpp, uses block-diffusion drafting to boost Qwen 3.6 27B inference speed, hitting 4.44x at 36K context with near-lossless quality.
Kernel 7.2: RK3588 media, smarter GPU memory management
Linux kernel 7.2 brings key improvements, including RK3588 support and Rust-based GPU memory management.
GPU-Tile-SIM: Tile-Centric GPU Simulation for LLM Hardware-Software Co-Design
GPU-Tile-SIM is a GPU simulation framework that offers high accuracy for LLM workloads in hardware-software co-design.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comDeepSeek-V4-Flash 0731: Full Precision Lossless Performance Test
DeepSeek-V4-Flash-0731 can be run at full precision with 176 GB memory. The project focuses on efficiency and performance testing.
672 GB VRAM with 7x RTX PRO 6000 Blackwell: More GPUs or 1 TB RAM?
Design for Kimi K3 project with 672 GB VRAM and 7x RTX PRO 6000 GPUs. Should we choose more GPUs or 1 TB of RAM?
AMD AITER MI3XX/CDNA3 Kernels Patched for MI2XX/CDNA2
AMD's AITER project provides performance enhancements for MI2xx and MI3xx cards.
Assessing LLM Inference Profitability: Insights from Kimi K3
Analysis and calculations on the profitability of LLM inference using Kimi K3.
Debugging Ray Tracing Applications with NVIDIA OptiX Toolkit
The NVIDIA OptiX Toolkit offers debugging tools for ray tracing applications, including error code checking and device-side debug printing.
AMD's CDNA5 Architecture Unveiled at Advancing AI 2026 Event
AMD unveiled its CDNA5 architecture at the Advancing AI 2026 event, focusing on efficiency gains and the transition to RDNA.
SkewAdam: A New Optimizer Reducing MoE State Memory by 97%
SkewAdam reduces MoE state memory by 97%, allowing a 6.7B model to fit on a 40GB GPU.