The problem

When does moving a page to where it’s used improve locality, and when does migration become the work?

What current research shows

CUDA exposes managed memory and a unified address model, while device performance guidance emphasizes minimizing host-device transfers and improving overlap. A page that moves across the interconnect adds latency and bandwidth demand.

Where the evidence stops

Migration behavior depends on access pattern, residency, prefetching, page size, interconnect, and concurrent users.

What Valen Systems is testing

Report page faults, migration volume, transfer time, residency, end-to-end work, and energy as workload behavior changes.

Sources

NVIDIA CUDA C++ Programming GuideNVIDIA CUDA C++ Best Practices Guide