» Tag
deep-learning
11 postsDCGAN Paper Shows Convolutional GANs Learn Reusable Image Features
DCGAN paper introduces a stable convolutional GAN architecture, showing learned features work as general-purpose image representations.
Audit Finds Most Distributional RL Risk Claims Are False
A new audit framework shows most risk claims from distributional RL agents like QR-DQN and C51 are training artifacts, not genuine environment risk signals.
FlowOptimizer: Learning to Optimize via Unfolded Flows
MIT and Boston University researchers unveil FlowOptimizer, a flow-based learning-to-optimize framework that outperforms classical and learned optimizers by orders of magnitude.
CommitBrief — AI code reviews, right in your terminal
A provider-agnostic, local-first CLI that reviews your staged changes, a historic range, or a whole GitHub pull request. Zero telemetry, no server. Free and open source.
commitbrief.comScaling Laws Explained: From Kaplan to Chinchilla to Overtraining
A breakdown of LLM scaling laws from Kaplan to Chinchilla, and why modern models are deliberately overtrained to cut inference costs.
MTIA 300: Meta's First Training Chip with Integrated NICs
MTIA 300 is Meta's first training chip optimized for recommendation models, featuring integrated NICs that revolutionize communication.
Mastering Edge AI: Building High-Speed Vision Analyzers on Android
Explore the balance between deep learning and mobile device constraints. Learn about the architecture and Kotlin patterns for high-speed vision analyzers.
Inkling by Thinking Machines: A Groundbreaking Multimodal LLM
Thinking Machines introduces Inkling, a multimodal LLM with 1 trillion parameters and 1M context window, offering new opportunities for engineers.
Reducing HBM Bottlenecks in JAX-Based LLM Training
Methods to enhance efficiency in LLM training by reducing HBM pressure with JAX.
DeepSWE: The Best Benchmark for Evaluating AI Coding Agents?
DeepSWE offers a novel benchmarking platform for evaluating the performance of AI coding agents.
SkewAdam: A New Optimizer Reducing MoE State Memory by 97%
SkewAdam reduces MoE state memory by 97%, allowing a 6.7B model to fit on a 40GB GPU.