The problem
Which work can run now, and which work must wait for a predecessor or transfer?
What current research shows
CUDA exposes streams, events, asynchronous operations, and execution graphs because GPU work isn’t a single flat queue. Dependencies constrain legal overlap, while independent work can expose concurrency.
Where the evidence stops
More concurrency can increase contention or memory pressure. A graph that’s theoretically parallel may not be practically efficient.
What Valen Systems is testing
Represent work as a dependency graph and measure overlap, idle gaps, synchronization delay, memory pressure, and completion time.