The problem

Where does useful work stop while one side waits for the other side's result or signal?

What current research shows

The CUDA programming model separates host code from device kernels and provides asynchronous streams and events to coordinate them. That boundary can expose synchronization overhead, transfer delay, dependency stalls, or insufficient work ahead of the device.

Where the evidence stops

The bottleneck can be CPU preparation, PCIe or interconnect transfer, GPU execution, memory, or synchronization itself.

What Valen Systems is testing

Measure end-to-end critical path, device idle time, host wait time, transfer volume, and overlap rather than optimizing one side in isolation.

Sources

NVIDIA CUDA C++ Programming GuideNVIDIA: writing CUDA kernels