Your new inference report shows a shorter wait for the first token. Before calling it a performance gain, check where the stopwatch stops. A benchmark migration can change that boundary without changing the machine at all.
For software companies serving AI workloads, platform teams and data-center operators, this matters when a chart becomes a capacity decision. An apparent improvement can support the wrong hardware purchase or hide an application that still keeps its users waiting. Start with comparable measurements, then decide what the system can usefully serve.
The same label can describe a different wait
NVIDIA's September 18, 2026 article introduces AIPerf as the successor to GenAI-Perf. Its migration guide identifies a specific comparison trap for reasoning models: the older tool waits for non-reasoning output when measuring time to first token. AIPerf counts the first token of either kind. Those numbers answer different questions.
The migration guide gives two useful mappings. Compare GenAI-Perf's time to first token with AIPerf's time to first output token, or TTFO. Compare the older output sequence length with AIPerf's Output Token Count, not its Output Sequence Length, which also includes reasoning. Preserve these definitions beside the results instead of merging reports by matching column names.
Choose the wait your application needs to control
AIPerf documents TTFO as the wait for the first non-reasoning output. Its token separation also depends on the response format: reasoning exposed in a separate field can be counted separately; reasoning embedded in ordinary content is included in the output count unless filtered. Check the endpoint and parser you actually use, not just the model's name.
Our recommendation is to write the user-facing question first. Are you trying to show an answer sooner, finish an agent action within a deadline, or serve more simultaneous sessions without unacceptable delays? Save a separate application-level check for that outcome. A first output token doesn't establish that a tool call was valid or that the final answer was useful.
Make the first comparison deliberately boring
Before testing a new server configuration, establish what changed in the measurement. Use an authorized test environment and a representative workload. The following is our proposed comparison record, not an upstream certification procedure.
- Pin the benchmark, model, tokenizer, serving runtime and hardware configuration. Keep the original reports and commands.
- Record the request mix, input and output lengths, concurrency or arrival rate, warmup and cache conditions. Explain any differences the two tools cannot reproduce exactly.
- Map each reported metric to its definition. Separate a changed measurement boundary from an observed change in execution.
- Check the responses as well as the timings. Keep correctness criteria and the accepted delay visible in the decision record.
- Only then vary the configuration under evaluation. Report the remaining limitations instead of presenting an uncontrolled comparison as a capacity gain.
Count the requests that meet the requirement
AIPerf's goodput measures requests per second that satisfy configured service-level thresholds. Its Good Request Fraction uses attempted requests as the denominator, including errors. That distinction helps prevent a run with dropped traffic from looking healthy merely because the surviving requests were fast.
For a buying decision, we recommend keeping that result beside tail latency, errors and the application checks. State the thresholds. State the offered load. Keep answer quality separate from latency compliance. A higher token rate alone doesn't tell the team how many customers it can serve acceptably, or whether more hardware is the next useful investment.
Turn the uncertainty into a testable decision
If you know the workload but haven't agreed what would justify a change, start there. Valen Systems' Compute Evaluation Plan defines one workload comparison, its controls, correctness checks and decision criteria. It's a planning deliverable, not an executed benchmark or a promised speedup. Running the evaluation needs its own scope. You can talk with us before choosing that starting point.
Sources
- NVIDIA: Benchmarking LLM Inference at Scale with AIPerfSource published September 18, 2026. Retrieved September 22, 2026.
- AIPerf: Migrating from GenAI-PerfPublication date not stated. Retrieved September 22, 2026.
- AIPerf: Metrics ReferencePublication date not stated. Retrieved September 22, 2026.
Put this to work.
Talk through the scope with Valen Systems
A smaller starting point: Plan a workload comparison