The problem

How many blocks can actually be active when every kernel consumes a different mix of scarce resources?

What current research shows

NVIDIA's Best Practices Guide describes occupancy as a result of execution parameters and per-block resource use, including registers and shared memory. More occupancy can help hide latency, but maximizing occupancy alone doesn’t guarantee maximum performance.

Where the evidence stops

Occupancy is a capacity indicator, not a universal throughput metric. A kernel can be limited by memory, instruction issue, dependencies, or a different resource.

What Valen Systems is testing

Record occupancy alongside achieved throughput, memory traffic, dependency stalls, power, and completion time for each kernel configuration.

Sources

NVIDIA CUDA C++ Best Practices GuideNVIDIA CUDA C++ Programming Guide