The problem

What happens when threads that were supposed to execute together choose different branches?

What current research shows

NVIDIA documents that branch divergence occurs within a warp and that different paths are executed serially for the affected threads. The result is lower effective utilization even though the hardware still exposes many threads.

Where the evidence stops

The impact depends on divergence frequency, path balance, compiler behavior, architecture, and whether divergence is hidden by other active warps.

What Valen Systems is testing

Measure branch efficiency, active warps, instruction throughput, and useful work across controlled workload shapes.

Sources

NVIDIA CUDA C++ Programming GuideNVIDIA CUDA C++ Best Practices Guide