The problem
Where does useful work stop while one side waits for the other side's result or signal?
What current research shows
The CUDA programming model separates host code from device kernels and provides asynchronous streams and events to coordinate them. That boundary can expose synchronization overhead, transfer delay, dependency stalls, or insufficient work ahead of the device.
Where the evidence stops
The bottleneck can be CPU preparation, PCIe or interconnect transfer, GPU execution, memory, or synchronization itself.
What Valen Systems is testing
Measure end-to-end critical path, device idle time, host wait time, transfer volume, and overlap rather than optimizing one side in isolation.