The problem
When is a second GPU useful after the cost of moving data and coordinating results?
What current research shows
The CUDA model exposes multiple devices and host-device transfers, while the performance guides emphasize minimizing transfers and overlapping transfer with computation where possible. A partition that balances arithmetic can still lose to communication cost.
Where the evidence stops
PCIe, NVLink, topology, memory capacity, collective operations, and workload shape determine the outcome.
What Valen Systems is testing
Report compute time, transfer time, link utilization, synchronization, scaling efficiency, and energy per completed workload.