Sunday, July 19, 2026

SemiAnalysis On Moonshot AI's Kimi K3: Probably Good For Nvidia and HBM; Not So Much For Open AI

Via Xitter: 

Thread at Thread Reader

Ending with a Jevons Paradox Cambrian Explosion, to mix a couple metaphors.

...More efficient attention will further push context lengths from 1M to 5M+, with less context rot. Jevons’ Paradox means that making attention more efficient will lead to wider AI adoption, which will require more networking. 8/8 

In essence the model is so big it will only be optimally run on Nvidia's premier offerings+High Bandwidth Memory+NVLink - Mellenox networks.

Here's their commentary:

[Emergency Episode] Moonshot’s Kimi K3 has Arrived! China has a Frontier Model