top of page
Search

A new paradAIgm

  • Writer: Gustavo A Cano, CFA, FRM
    Gustavo A Cano, CFA, FRM
  • 2 days ago
  • 2 min read

There is a new sheriff in town, and it comes from the East. Kimi-k3, it’s a new Chinese open model that is challenging the frontier (closed) models and is currently ranking number 1 as you can see on the top chart below. What are the implications? (1) K3 is open-weight, which means enterprises can run it on-premises or through third-party inference providers. This is a direct threat to the hyperscaler lock-in model, because If you can get Claude/Opus-level performance at a fraction of the cost. without being locked into Azure or AWS, the cloud arbitrage game changes completely. (2) K3 can hold entire codebases or book-length documents in a single prompt, which is incredibly more efficient. That means less cost, and that means compression of margins for OpenAI or Anthropic at a point where their valuations are severely stretched. (3) The token cost-per-intelligence curve is collapsing. K3 uses a combination-of-Experts architecture (swarm) where only 16 of 896 experts fire per token​. That means you're getting near-frontier reasoning at a fraction of the compute cost. You can see that in the lower chart below. Kimi-k3 is the only model within the green square, which is where you want your model to be. What’s the bottom line? If token costs keep falling while capability keeps rising, the companies that win won't be the ones selling the most GPUs; they'll be the ones building the most token-efficient platforms. That's why NVIDIA's inference-optimized chips matter as much as their training GPUs, and why hyperscalers are racing to build their own chips. Furthermore, there are second order effects: (1) Power constraints are the new chip shortages. If models like K3 deliver more intelligence per watt, the hyperscaler with the best efficiency strategy wins, not just the most GPUs. (2) Enterprise AI spend is pulling forward cloud migration, but also pulling it away from incumbents. With open-weight models and competitive API pricing, they're asking: Do we need AWS, or do we need compute? The race continues but the horses that came out of the box first, are being passed by a new type that’s more efficient and almost as powerful. And it will reshape several business models.


Want to know more? You can register for free at Fund@mental.




 
 
 

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page