A new memory architecture gets announced. Before it becomes a procurement requirement, someone should be able to finish this sentence: our workload spends its time waiting for ___. If the blank is still empty, the next purchase decision needs a measurement plan, not another specification comparison.

NVIDIA's NVHBM announcement is a useful prompt for that conversation. This article separates what NVIDIA announced from Valen Systems' proposed way to evaluate a workload. We haven't benchmarked NVHBM.

What NVIDIA announced

On August 26, 2026, NVIDIA announced NVHBM custom high-bandwidth memory as an expansion of NVLink Fusion for semi-custom AI infrastructure. Its described design moves the memory controller from the compute die into the HBM base die. NVIDIA presents that arrangement as a way to improve memory performance and efficiency while freeing compute-die space.

Those are the vendor's architectural description and positioning. They aren't Valen Systems measurements, a promise about an existing GPU, or evidence that changing memory will improve your application.

Source: NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory

A faster component isn't an application result

NVIDIA's performance guide distinguishes work limited by memory bandwidth, mathematical throughput, or latency. It uses arithmetic intensity—the amount of computation relative to data movement—to explain why the same hardware doesn't constrain every operation in the same way. The guide also warns that its simplified analysis doesn't replace more accurate profiling.

That's a useful technical starting point. Our practical recommendation is to identify the waiting in the workload you actually intend to run before deciding which hardware characteristic deserves more money.

Source: NVIDIA: GPU Performance Background User's Guide

Write the useful-output test first

Specify what a completed unit of useful work means. For a training job, that might include an agreed quality target and reproducible input conditions. For inference, it might include a defined request set, output checks, and a latency requirement. Pick the definitions appropriate to the application; don't substitute a convenient hardware counter for the result the business needs.

Record the software versions, model or application configuration, precision, input sizes, and concurrency assumptions. Decide which of those are fixed in the comparison and which are deliberately being changed. If several change at once, label the result as a system comparison rather than claiming one component caused the difference.

Build an evaluation around six decisions

These are our proposed planning questions for a compute evaluation, not test results or a purchasing recommendation for NVHBM.

  • Baseline: which current configuration and representative workload will the alternative be compared against, and why is that comparison useful?
  • Waiting: which observations will distinguish application work from queuing, data preparation, transfers, synchronization, or other delay?
  • Correctness: what output checks must pass before a faster run can count as an improvement?
  • Repeatability: how will warm-up, run order, background activity, and repeated trials be recorded so a surprising result can be investigated?
  • Economics: which costs belong in the decision, including migration work and operating constraints rather than only a component price?
  • Decision threshold: what result would justify moving forward, and what result would cause the team to keep the existing system?

Keep the measurement boundary visible

If energy is part of the evaluation, state what is being measured: a reported device value, a server, or an agreed facility boundary. Don't relabel one as another. Record which equipment and supporting work sit outside that boundary. Have the responsible facilities team approve any site-level measurement approach.

Do the same for time. A faster execution segment doesn't automatically shorten an end-to-end job by the same amount. Keep the segment result and the user-visible completion result separate in the report. If a required measurement isn't available, mark the conclusion as unresolved instead of filling the gap with a vendor headline.

Leave room for a result that doesn't sell new hardware

An evaluation should be allowed to conclude that the present workload doesn't justify a migration, or that software and operating changes need investigation first. Define that possibility before testing. Otherwise the exercise risks becoming a search for evidence to support a purchase already decided.

Valen Systems' Compute Evaluation Plan scopes one workload and up to two configurations into a measurement plan. It doesn't include an executed benchmark, hardware access, or a guaranteed efficiency improvement. Bring the workload, the current constraint, and the decision you need to make; the first deliverable is a testable question and a clear plan for answering it.

Sources

  1. NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth MemorySource published August 26, 2026. Retrieved September 9, 2026.
  2. NVIDIA: GPU Performance Background User's GuidePublication date not stated. Retrieved September 9, 2026.

Put this to work.

Plan a workload evaluation

Have a different situation? Talk with Valen Systems before choosing a scope.