The world's largest open model goes rentable: Kimi K3 lands on CoreWeave
Kimi K3, at 2.8 trillion parameters with a 1M-token context, now runs on CoreWeave Dedicated Inference. Bring your own weights and get an autoscaled endpoint on NVIDIA GPUs.
The flagship of the open-source camp has crossed another threshold: CoreWeave now serves Kimi K3, at 2.8 trillion parameters with a 1-million-token context window, on its Dedicated Inference service. The company's announcement bills it as "the world's largest open-source model", and the flow is reduced to three steps: upload, create a gateway, go.
The part that matters to developers is the flexibility: you can bring your own weights. If you have fine-tuned Kimi K3, you can load that custom version onto CoreWeave's NVIDIA GPUs and have it served with autoscaling and routing. That turns the real obstacle in front of giant open models, the question of who runs this thing, where and how, into something you can simply rent.
For anyone following Kimi K3's journey, the picture is filling in: Applied Compute opened the full training recipe last week, and now the hosting side has hit the commercial shelf. Training open, weights open, and operating it now a click away: the distance between a 2.8T open model and the closed labs' price lists narrows a little more each week.
One field note is due: Dedicated Inference is invite-only for now, and that drew fire on day one. Developer Ed Sealing, who says he is a CoreWeave shareholder, replied under the announcement that he sent a corporate invite request months ago and never heard back. The technology behind the door is compelling; the door itself opens a little slowly.