The problem

Which work can run now, and which work must wait for a predecessor or transfer?

What current research shows

CUDA exposes streams, events, asynchronous operations, and execution graphs because GPU work isn’t a single flat queue. Dependencies constrain legal overlap, while independent work can expose concurrency.

Where the evidence stops

More concurrency can increase contention or memory pressure. A graph that’s theoretically parallel may not be practically efficient.

What Valen Systems is testing

Represent work as a dependency graph and measure overlap, idle gaps, synchronization delay, memory pressure, and completion time.

Sources

NVIDIA CUDA C++ Programming GuideNVIDIA CUDA C++ Best Practices Guide