Applied Compute opens full Kimi K3 training: GPUs per replica down 40 percent

Fine-tuning Moonshot's 3-trillion-parameter open-weight model on AC2: custom activation operators free ~90 GB per GPU, and MXFP4 inference runs on 2 nodes instead of 4.

Paylaş
Applied Compute opens full Kimi K3 training: GPUs per replica down 40 percent

Applied Compute has opened full fine-tuning of Kimi K3 - Moonshot's 3-trillion-parameter open-weight model - on its AC2 agent cloud, and published the engineering report behind it. The company's memory work cuts the GPUs needed per training replica by roughly 40 percent - a meaningful cost threshold for training open models at frontier scale.

Most of the gain comes from three places. Custom operators for the SiTU-GLU activations recompute intermediate tensors during backpropagation instead of storing them, freeing about 90 GB of high-bandwidth memory per GPU - roughly a third of total GPU memory. Chunked Adam optimizer steps shave 4 of the original 12 bytes per parameter off host memory. On the inference side, MXFP4 low-precision rollout engines run on 2 B300 nodes where bf16 needed 4; the measured KL divergence is comparable to bf16's own training-inference mismatch, in part because only 47 percent of active parameters - about 49 of 104 billion - drop to low precision.

The report also covers the less glamorous work: online weight transfers were found to corrupt rollouts through the Block Attention Residuals cache, now recalculated on every update; and writing a 28-terabyte checkpoint sped up 2.2x on Weka and 4.4x on NFS once Linux page caches were bypassed with O_DIRECT.

The company puts the result concretely: after 50 training steps within a single day, a coding agent matched performance that a smaller 300-billion-parameter model needed multiple days to reach. The improvements are not specific to Kimi K3, it says - they carry over to the platform's other models as well.

Source: Applied Compute report