» Tag
gpu
91 postsLarge-Scale TensorCircuit Contractions: Disabling XLA GPU Autotuning
Impact of disabling XLA GPU autotuning on memory savings and runtime for large TensorCircuit contractions.
Llm.c Ported to Mojo with Metal Kernels: 1.72x Faster than PyTorch MPS
The Mojo port of llm.c with Metal kernels achieves 1.72x faster performance than PyTorch MPS.
CTA-Pipelining: A Latency-Oriented Scaling Method for Multi-GPU Systems
CTA-pipelining enhances performance in multi-GPU systems by focusing on latency-oriented scaling.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comReducing HBM Bottlenecks in JAX-Based LLM Training
Methods to enhance efficiency in LLM training by reducing HBM pressure with JAX.
Understanding NVIDIA DGX Spark Environment: From API to GPU, Week 1
Get essential information and commands to start working in the NVIDIA DGX Spark environment. Learn the details of the first week.
Catch your local LLM falling back to CPU
Picchio is a Python file that checks your local LLM setup, showing if the GPU really did the work.
KubeSwift: Kubernetes-Native VM Orchestration on Cloud Hypervisor
KubeSwift defines VMs as Kubernetes CRDs and runs them on Cloud Hypervisor with GPU passthrough, live migration, and OCI image distribution.
DFlash in llama.cpp: 4.44x Faster Local Inference on Qwen 3.6 27B
DFlash, merged into llama.cpp, uses block-diffusion drafting to boost Qwen 3.6 27B inference speed, hitting 4.44x at 36K context with near-lossless quality.
Kernel 7.2: RK3588 media, smarter GPU memory management
Linux kernel 7.2 brings key improvements, including RK3588 support and Rust-based GPU memory management.
GPU-Tile-SIM: Tile-Centric GPU Simulation for LLM Hardware-Software Co-Design
GPU-Tile-SIM is a GPU simulation framework that offers high accuracy for LLM workloads in hardware-software co-design.