Kimi K3: A Scaling Experiment Beyond the 2.8T Model
Kimi K3 presents a scaling experiment that surpasses the 2.8T model. Insights on expert routing and load balancing are discussed.
The Kimi K3's 2.8 trillion parameters initially attract attention, but the real intrigue lies in how the model scales capacity while controlling active computation. K3 activates only 16 of 896 experts per token, indicating a vast storage capacity with minimal participation in each forward pass. The economics of long context also shift; at this scale, it becomes a memory and serving issue rather than a capability problem.