The problem
How much of each fetched memory transaction does the warp actually use?
What current research shows
NVIDIA's current kernel-writing guide explains that a warp combines thread requests into the transactions needed to satisfy the addresses, and that coalesced access uses the memory system more efficiently. Scattered access requires more transactions and transfers more unused data.
Where the evidence stops
The exact transaction behavior depends on architecture, word size, alignment, cache state, and access pattern.
What Valen Systems is testing
Compare memory transactions, achieved bandwidth, cache behavior, and work per joule across contiguous and scattered access patterns.