Unlocking asynchronicity in continuous batching for LLM inference

We are leaving serious performance on the table when LLM inference loops force our CPUs and GPUs to take turns....

Feed
October 4, 2026
Unlocking asynchronicity in continuous batching for LLM inference


Hardware costs money. Renting an H200 runs about five dollars an hour, which balloons past a hundred bucks a day if you leave it spinning. When you're footing that kind of bill, leaving hardware idle feels like burning cash. We talk endlessly about optimizing model weights, quantization, and squeezing every ounce of efficiency out of our attention ways. Yet, a massive bottleneck hides in plain sight right within our orchestration loops. Continuous batching solved padding waste by keeping requests tightly packed, but it missed a fundamental flaw in how the control flow operates.

By default, standard batching is stubbornly synchronous. The CPU and the GPU take turns waiting for each other like polite strangers at a narrow doorway. While the GPU crunches a forward pass, the CPU sits entirely idle. The moment that compute finishes, roles reverse. The CPU takes over to sample tokens, update KV cache tables, evict finished sequences, and stage the next batch while the expensive silicon sits cold. Running hundreds of inference steps every single second turns these micro-delays into a catastrophic tax. Those idle gaps quickly consume nearly a quarter of your total runtime.

Unlocking asynchronicity in continuous batching for LLM inference

Fixing this requires real architectural separation. We need to completely decouple CPU batch preparation from GPU execution. If the CPU prepares the next iteration ahead of time while the GPU is still crunching the current one, everything changes. The hardware stays pegged at maximum utilization. No more alternating green and red blocks on your trace timelines. Just relentless, overlapping throughput that actually justifies the cloud bill.

Good engineering isn't just about picking the right model. It demands that we look closely at the invisible orchestration taxes we pay on every single token. Stop letting your hardware wait around.