» Tag
gpu
91 postsConfidential Computing: Protecting Data in Use on CPU and GPU Systems
Confidential computing enhances data protection in AI data centers. It safeguards sensitive information on CPU and GPU systems.
Ninfer: High-performance single-GPU inference
NInfer offers a high-performance C++/CUDA inference engine for RTX 5090.
RL-Training Agent Developed for Model Training
Learn about the new pipeline developed for model training using an AI agent.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comEnki: Write GPU Compute Kernels in Pure Stable Rust
Enki enables writing GPU compute kernels in Rust, allowing dynamic execution on CPU and GPU for enhanced performance.
Testing LLM Concurrency on Consumer Hardware (RTX 5060)
LLM concurrency tests on RTX 5060 yield crucial insights for engineers.
Co-Designing AI Model Attention for Fast, Interactive Long-Context Inference
Explore practical guidelines for optimizing AI model attention and inference efficiency.
New Storage Technology Could Enable GPUs to Reach Terabyte Capacities
High-Bandwidth Flash (HBF) could boost GPU memory capacity to terabytes, enhancing AI systems. Discover the implications of this new technology.
GLM-4.7-Flash on 2x RTX 3090: My Hands-On Experience
GLM-4.7-Flash was tested on 2x RTX 3090. Performance comparison in short and long contexts was conducted.
Voice Chat Server on RTX 3050 Ti: Achieving 11.9s Voice-to-Voice
Learn about the voice chat server setup on RTX 3050 Ti and its 11.9s response time.
Daemon Automatically Switches GPU Mode and Refresh Rate for Linux Laptops
A new daemon for Linux laptops saves battery by automatically managing GPU mode and refresh rate based on power state.