On September 23, Inferact published a Kimi K3 inference benchmark alongside the Apache 2.0 repository inferact/tpu-megakernels. The startup was founded by core members of the vLLM team and had raised a $150 million seed round at an $800 million valuation, led by a16z and Lightspeed. Its test compared two 16-chip systems running Moonshot AI’s 92-layer mixture-of-experts model, which combines Kimi Delta Attention with multi-head latent attention (MLA). The chips used 4-way tensor parallelism and 8-way expert parallelism.

On paper, GB200 has the stronger compute and bandwidth specifications. Each chip is reported to deliver 2.5 PFLOPS of BF16 and 5 PFLOPS of FP8, with 186GB of HBM at 8000 GB/s. TPU v7 Ironwood is reported at 2.31 PFLOPS BF16 and 4.61 PFLOPS FP8, with 206GB of HBM at 7380 GB/s. TPU v7’s hardware edge is 20GB more memory per chip. Yet Inferact reported 709 tokens per second on TPU v7, against 452 on GB200, a roughly 57% difference.
But the test changed the software as well as the chip.

Inferact ran its megakernel on TPU v7 and vLLM on GB200. The result therefore measures two hardware-and-software combinations, not an isolated silicon contest. It shows that Inferact’s TPU implementation outperformed the GB200 control on this workload. It does not establish that TPU v7 is inherently faster, or that the same gap would survive equivalent tuning on Nvidia.
Inferact’s megakernel targets work between tokens

A megakernel combines operations that would otherwise run in separate kernels. In general, this approach can reduce launch overhead and the movement of intermediate data. That can matter in autoregressive decoding, where a model performs another round of work for each generated token, especially when a small batch leaves less work to run concurrently.
Kimi K3’s MoE routing and mixed attention architecture make its execution path a particular workload, but the published figures alone do not reveal which operations Inferact fused or how much each contributed. The safe conclusion is narrower: a specialized TPU implementation posted higher reported throughput than the vLLM GB200 control in this configuration. The mechanism is credible; its precise contribution remains unquantified in the source report.
The authorship makes the comparison notable. Inferact was founded by core members of the team behind vLLM, which ran on the GB200 side. That lends context to the choice of baseline, but it does not make the implementations equivalent. The reported win is for Inferact’s specialized stack over vLLM on this test.
Batch-size results add useful context. Without speculative decoding, TPU v7 recorded 249 tokens per second at batch size 1, versus 127 for GB200; at batch 2, 392 versus 227; at batch 4, 515 versus 373; and at batch 8, 865 versus 636. The relative gap is largest at batch 1 and smaller at batch 8. That pattern is consistent with overhead playing a greater role at smaller batches, though the figures do not isolate launch overhead from other differences between the two implementations.
The 57% result has clear limits
The report gives a specific setup: 16 chips per system, Kimi K3, 4-way tensor parallelism, 8-way expert parallelism, and single-stream decoding for the headline comparison. It also says the benchmark and Apache 2.0 code appeared on September 23. Inferact’s reported funding and valuation provide company context, not independent confirmation of the measurements.
The source is a Cocoloop Editorial report localized from public sources. The figures should be treated as reported, not independently audited. The benchmark does not establish results for other models, parallelism configurations, or serving conditions. Nor does it compare cost, power consumption, or total cost of ownership.
The main qualification is the control: GB200 ran vLLM, while TPU v7 ran Inferact’s megakernel. The report provides no equivalent Nvidia megakernel result. This is evidence of a stack-level performance difference on one workload, not a general ranking of TPU v7 and GB200.
A Nvidia port could erase TPU v7’s reported lead
The most important implication depends on portability. Kernel fusion and overhead reduction are software techniques, but the benchmark does not show that Inferact’s implementation can move unchanged to Nvidia hardware. Porting may require substantial engineering, and CUDA’s programming model, hardware behavior, and surrounding runtime all matter. A technique can be portable in principle without delivering the same gain on another accelerator.
My forecast is that Inferact or a competitor will demonstrate a megakernel approach on Nvidia hardware within 12–24 months, and that optimized kernel layers will become a standard part of frontier inference. That call rests on an incentive as much as a technical possibility: a 57% reported gap on a prominent workload gives competing teams reason to test whether better kernel design can recover performance on GPUs too. If a serious Nvidia implementation closes much of the gap, the current TPU result will look less like a durable hardware advantage and more like a lead created by software maturity.
Near parity is a possibility, not a consequence guaranteed by the published peak specifications. GB200 leads the reported per-chip compute and bandwidth figures, but those figures alone cannot predict realized decode speed. A Nvidia port could narrow the observed difference; it could also fail to reproduce the TPU gains, or expose hardware-specific advantages that preserve a gap. I would revise the forecast if comparable Nvidia tuning failed to approach the TPU result, or if the megakernel’s performance depended on TPU-specific execution features.
If the approach does travel across platforms, the competitive prize shifts toward whoever can build and maintain the optimization layer. Inferact’s open-source release may help establish that layer, but the benchmark does not prove that the company will become a commercial toll booth or a default vendor. Its reported $800 million valuation reflects investor expectations; this one result is not enough to validate them. The business case would depend on sustained performance gains, support across more models and chips, and a way to capture value from software released under Apache 2.0.
For Nvidia, the result is a reminder that strong hardware specifications do not guarantee the best observed throughput when software differs. For Google, it is a credibility point with an adoption condition: the reported result depends on Inferact’s stack, not TPU silicon alone. Enterprise buyers comparing the two systems need benchmark results that hold the model, software maturity, serving conditions, cost, and power in view. The Cocoloop report does not supply that full comparison.
The September 23 result reverses the paper comparison: GB200 leads on the reported compute and bandwidth figures, while TPU v7 leads in this particular decoding test. The unresolved test is whether a comparable megakernel can reproduce the gain on Nvidia. That result, not the 57% headline by itself, will show whether Inferact has built a portable performance layer or a strong TPU-specific implementation.