» Tag
gpu
91 postsBenchmarking Inference Energy Costs of LLM: A LLaMA Study
Researchers benchmark LLaMA model sizes on V100 and A100 GPUs to analyze the energy and compute costs of LLM inference at scale.
NoGraphicsAPI Reimagines Vulkan Without Descriptor Sets or Pipeline Layouts
NoGraphicsAPI is an experimental Vulkan 1.4 backend replacing descriptor sets and pipeline layouts with GPU pointers and app-owned heaps.
Inside NVIDIA Exemplar Cloud: Closing 8-12% AI Training Performance Gaps
How NVIDIA diagnosed 8-12% AI training performance gaps on H100 and GB200 NVL72 clusters using perf, Nsight Systems, SMMU tuning, and NUMA fixes.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comSkyPilot and Hugging Face bring zero-egress storage to multi-cloud AI
SkyPilot and Hugging Face add a zero-egress storage backend, letting AI teams run GPU workloads on any cloud while data stays on the Hub.
Speculative Decoding: The Free Speedup Most Local LLM Setups Skip
Speculative decoding speeds up local LLM inference 1.5-2.5x with identical output; 2026 saw it built into models via multi-token prediction.
ZML/LLMD Alpha: One LLM Server Across CUDA, ROCm, TPU, Metal
ZML/LLMD alpha runs LLaMa, Gemma, Qwen and Mistral models across NVIDIA, AMD, TPU, Intel and Apple Metal in one server, with DFlash speeding up inference.
Restore LLM Inference Capacity Quickly with Shadow Engine Recovery in NVIDIA Dynamo
NVIDIA Dynamo's shadow engine recovery allows rapid restoration of LLM processes after failures.
The Cost of Irregularity: CUDA C++, Rust, and Triton
A comparison of CUDA C++, Rust, and Triton in GPU programming. Performance differences in irregular workloads are analyzed.
From CUDA to MLX: K-Search Optimizes for Apple Silicon
K-Search translates CUDA optimizations into MLX for Apple Silicon, marking a key advancement for GPU kernels in AI.
XY: A Fast and Customizable GPU-Accelerated Plotting Library
XY is a fast, interactive Python charting library accelerated by GPU.