» Tag
gpu
82 postsGLM-4.7-Flash on 2x RTX 3090: My Hands-On Experience
GLM-4.7-Flash was tested on 2x RTX 3090. Performance comparison in short and long contexts was conducted.
Voice Chat Server on RTX 3050 Ti: Achieving 11.9s Voice-to-Voice
Learn about the voice chat server setup on RTX 3050 Ti and its 11.9s response time.
Daemon Automatically Switches GPU Mode and Refresh Rate for Linux Laptops
A new daemon for Linux laptops saves battery by automatically managing GPU mode and refresh rate based on power state.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comLarge-Scale TensorCircuit Contractions: Disabling XLA GPU Autotuning
Impact of disabling XLA GPU autotuning on memory savings and runtime for large TensorCircuit contractions.
Llm.c Ported to Mojo with Metal Kernels: 1.72x Faster than PyTorch MPS
The Mojo port of llm.c with Metal kernels achieves 1.72x faster performance than PyTorch MPS.
CTA-Pipelining: A Latency-Oriented Scaling Method for Multi-GPU Systems
CTA-pipelining enhances performance in multi-GPU systems by focusing on latency-oriented scaling.
Reducing HBM Bottlenecks in JAX-Based LLM Training
Methods to enhance efficiency in LLM training by reducing HBM pressure with JAX.
Understanding NVIDIA DGX Spark Environment: From API to GPU, Week 1
Get essential information and commands to start working in the NVIDIA DGX Spark environment. Learn the details of the first week.
Catch your local LLM falling back to CPU
Picchio is a Python file that checks your local LLM setup, showing if the GPU really did the work.
KubeSwift: Kubernetes-Native VM Orchestration on Cloud Hypervisor
KubeSwift defines VMs as Kubernetes CRDs and runs them on Cloud Hypervisor with GPU passthrough, live migration, and OCI image distribution.