The problem
Which shared resource becomes the bottleneck when multiple kernels overlap?
What current research shows
NVIDIA's performance guidance treats global memory, caches, shared memory, registers, and execution configuration as coupled performance constraints. A kernel that fits in one resource can still contend in another.
Where the evidence stops
Hardware counters and profiler definitions differ by architecture. A single occupancy number can’t identify every contention source.
What Valen Systems is testing
Use counters and controlled co-runs to separate memory, cache, execution, copy, and synchronization contention.