We're open-sourcing Mixture-of-Kittens (MoK), our MoE training megakernel for NVL72s.
It fuses all Mixture-of-Experts communication and computation into a single, fully deterministic kernel, and runs up to 2.37x faster than the strongest public baselines.
Cursor on X: "We're open-sourcing Mixture-of-Kittens (MoK), our MoE training megakernel for NVL72s. It fuses all Mixture-of-Experts communication and computation into a single, fully deterministic kernel, and runs up to 2.37x faster than the strongest public baselines. https://t.co/yHu5E6RXp9"
- MoK now powers training across tens of thousands of GPUs at Cursor. In production, it raised end-to-end training throughput by 1.41x over our previous DeepEP-based stack.Our hope is that this lowers the barrier to AI research, so more labs can train models efficiently. Here's how we built it:

