The problem
When does combining work improve device utilization, and when does it delay the first useful result?
What current research shows
The CUDA programming model launches kernels from host code and allows asynchronous work. Many small launches expose more host-device coordination points; larger kernels or graphs can reduce launch and synchronization overhead but may increase latency or reduce flexibility.
Where the evidence stops
The best granularity depends on launch path, workload size, dependencies, device, and latency target.
What Valen Systems is testing
Measure time to first result, total completion time, launch count, synchronization points, and energy across batching choices.