Via Xitter:
Similar to DeepSeek in January 2025, Panicans may think that the AI networking switch TAM will massively shrink because Kimi K3 uses KDA Attention, which reduces KV-transfer networking bandwidth by up to 10x. But the opposite is true, as we explain below. 👇️ 1/8🧵 pic.twitter.com/FNrFeHzRWi
— SemiAnalysis (@SemiAnalysis_) July 18, 2026
Ending with a Jevons Paradox Cambrian Explosion, to mix a couple metaphors.
...More efficient attention will further push context lengths from 1M to 5M+, with less context rot. Jevons’ Paradox means that making attention more efficient will lead to wider AI adoption, which will require more networking. 8/8
In essence the model is so big it will only be optimally run on Nvidia's premier offerings+High Bandwidth Memory+NVLink - Mellenox networks.
Here's their commentary:
[Emergency Episode] Moonshot’s Kimi K3 has Arrived! China has a Frontier Model
Kimi K3 2.8T is so large that it will not fit on a single NVIDIA DGX B200, even at FP4. A GB300 NVL72, B300, or MI355X system is required, as each GPU has 288 GB of memory.
— SemiAnalysis (@SemiAnalysis_) July 17, 2026
One optimization that could make Kimi K3 fit on B200 is to gang multiple nodes together and use a… pic.twitter.com/7Yu1dJoPm0