» Tag
cuda
16 postsThe Cost of Irregularity: CUDA C++, Rust, and Triton
A comparison of CUDA C++, Rust, and Triton in GPU programming. Performance differences in irregular workloads are analyzed.
From CUDA to MLX: K-Search Optimizes for Apple Silicon
K-Search translates CUDA optimizations into MLX for Apple Silicon, marking a key advancement for GPU kernels in AI.
WISP: A CUDA Engine for Streaming 744B+ Parameter MoE Models on Consumer Hardware
WISP is a CUDA engine for streaming 744B+ parameter MoE models on consumer hardware.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comNinfer: High-performance single-GPU inference
NInfer offers a high-performance C++/CUDA inference engine for RTX 5090.
Anatomy of a CUDA Binary
Explore the structure of CUDA binaries and EIATTR encoding. Learn about Nvidia's undocumented features.
Raising the baseline for the `nvptx64-nvidia-cuda` target
Rust 1.97 increases the PTX ISA and GPU architecture for `nvptx64-nvidia-cuda`, making older driver compatibility impossible.