The problem

When does combining work improve device utilization, and when does it delay the first useful result?

What current research shows

The CUDA programming model launches kernels from host code and allows asynchronous work. Many small launches expose more host-device coordination points; larger kernels or graphs can reduce launch and synchronization overhead but may increase latency or reduce flexibility.

Where the evidence stops

The best granularity depends on launch path, workload size, dependencies, device, and latency target.

What Valen Systems is testing

Measure time to first result, total completion time, launch count, synchronization points, and energy across batching choices.

Sources

NVIDIA CUDA C++ Programming GuideNVIDIA CUDA C++ Best Practices Guide