» Tag
gpu
82 postsUnified Memory: Why Mini PCs Run 70B Models a Big GPU Can't
Unified-memory mini PCs like AMD's Strix Halo can load 70B-parameter models that a $2,000 RTX 5090 cannot fit. Here's why capacity and bandwidth pull in opposite directions.
The First Open-Source Agent Skills Collection for AMD ROCm
While NVIDIA has 428+ agent skills on skills.sh, AMD had none. This new open-source project delivers 10 production-ready skills for ROCm GPU workflows.
The AI Data-Centre Bust Will Look Like a Boom
The growing demand for AI boosts the importance of custom chips, but GPU purchases with debt could trigger a bust in data centers.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comMaking Long-Running NVIDIA TensorRT Engine Builds Observable and Cancelable
Learn to use the IProgressMonitor API for observable and cancelable TensorRT engine builds.
Confidential Computing: Protecting Data in Use on CPU and GPU Systems
Confidential computing enhances data protection in AI data centers. It safeguards sensitive information on CPU and GPU systems.
Ninfer: High-performance single-GPU inference
NInfer offers a high-performance C++/CUDA inference engine for RTX 5090.
RL-Training Agent Developed for Model Training
Learn about the new pipeline developed for model training using an AI agent.
Testing LLM Concurrency on Consumer Hardware (RTX 5060)
LLM concurrency tests on RTX 5060 yield crucial insights for engineers.
Co-Designing AI Model Attention for Fast, Interactive Long-Context Inference
Explore practical guidelines for optimizing AI model attention and inference efficiency.
New Storage Technology Could Enable GPUs to Reach Terabyte Capacities
High-Bandwidth Flash (HBF) could boost GPU memory capacity to terabytes, enhancing AI systems. Discover the implications of this new technology.