A service that stops crashing but takes too long to answer still has an operating problem. Before treating a memory-policy change as recovered capacity, measure whether the workload completes correct work within its required response time. That gives a data-center team something useful to compare with buying more hardware.

Kubernetes' September 14 announcement makes that test timely. Memory QoS is beta in v1.37, with its feature gate enabled by default on the kubelet. The default configuration still leaves memory throttling and reservation inactive. Check the settings actually applied to your nodes before attributing a result to the upgrade.

Source: Kubernetes v1.37: Memory QoS Graduates to Beta

Read the effective policy before changing it

For Linux nodes using cgroup v2, the announcement identifies two controls. memoryThrottlingFactor now defaults to null; an explicit value enables memory.high throttling. memoryReservationPolicy defaults to None; TieredReservation enables tiered protection. An existing explicit throttling factor survives the upgrade, but a configuration that relied on the former default needs attention.

For an evaluation, record the kubelet version and effective configuration beside the resulting cgroup values. Keep that record with the workload inputs. A comparison named only 'before upgrade' and 'after upgrade' won't tell the next operator which policy was tested.

Source: Kubernetes v1.37: Memory QoS Graduates to Beta

A running process can still be losing useful time

The Linux kernel documents memory.high as a throttling boundary: crossing it puts processes under reclaim pressure, rather than directly invoking the out-of-memory killer. A workload can therefore remain alive while its performance deteriorates. memory.max is the separate hard-limit mechanism. Fewer restarts alone won't establish that the change helped.

Use memory.events to distinguish high-boundary throttling from out-of-memory events, and be explicit about whether you're reading hierarchical or local counters. Pressure Stall Information, or PSI, adds a different observation: time tasks spend stalled on a resource. Linux exposes memory pressure at system level and, with the required support, per cgroup.

Correlate those signals with the work the service is meant to do. Did valid requests finish on time? Did a batch finish within its window? Did a neighboring workload lose responsiveness? Those are the acceptance questions; a calmer restart chart is supporting evidence.

Source: Linux kernel: Control Group v2, memory controller reference Source: Linux kernel: PSI, Pressure Stall Information

Include the neighboring workload in the test

The Kubernetes announcement flags a node-wide limitation: the reservation policy applies across its Pods, without individual opt-out. Protected memory can also include page cache. That makes the mix of work sharing a node relevant to the evaluation.

Our recommendation is to test a representative contention case in a controlled environment, with an agreed stopping point. An isolated run may answer whether one process finishes. It won't answer whether the proposed policy is acceptable while the other work you intend to colocate is active.

Source: Kubernetes v1.37: Memory QoS Graduates to Beta

Write a small comparison that can change a decision

Start with one workload and one proposed policy change. Hold the inputs and correctness checks constant. Write down the acceptable response time or completion window before looking at the results. For a data-center capacity decision, we'd want the evaluation plan to answer these questions.

  • Baseline: what exact software, node configuration, workload inputs and colocated work produced the current result?
  • Change: which memory setting will differ, and how will you verify it took effect?
  • Outcome: how much correct work finishes within the agreed window, including slow responses and failed attempts?
  • Cost to neighbors: what happens to their latency, completion times, memory pressure and failures under the same load?
  • Repeatability: can you repeat both cases, explain warm-up and cache conditions, and separate a persistent effect from run-to-run noise?
  • Decision: what result justifies a wider trial, what triggers rollback, and who accepts the tradeoff?

Make the next purchase answer that question

If the evidence points to insufficient capacity, it gives you a better hardware requirement. If a policy change meets the same workload target with acceptable effects on its neighbors, it gives you a candidate for a wider trial. Neither conclusion should be written in advance.

Valen Systems' Compute Evaluation Plan turns one workload question and up to two proposed configurations or approaches into a written measurement plan, correctness checks and decision criteria. It's a planning deliverable, not an executed benchmark or a promised performance improvement. Use the review below, or talk with us first if the question spans several workloads or sites.

Sources

  1. Kubernetes v1.37: Memory QoS Graduates to BetaSource published September 14, 2026. Retrieved September 18, 2026.
  2. Linux kernel: Control Group v2, memory controller referencePublication date not stated. Retrieved September 18, 2026.
  3. Linux kernel: PSI, Pressure Stall InformationPublication date not stated. Retrieved September 18, 2026.

Put this to work.

Plan a workload comparison

Have a different situation? Talk with Valen Systems before choosing a scope.