A cluster disappears from the control plane. The replacement workload starts somewhere else. Can a customer still complete the transaction that matters? That's the question a recovery review needs to answer, beyond whether a scheduler placed the workload.
Karmada's graduation is a reason to look at the project. It isn't a substitute for evidence about your own application. The distinction below is Valen Systems' analysis: project maturity and operational readiness are separate decisions.
What changed at the project level
CNCF has announced Karmada's graduation. Karmada is an open-source system for orchestrating Kubernetes workloads across multiple clusters. The announcement describes centralized placement, workload propagation, failover, and scaling, alongside graduation work on security and governance.
The announcement's URL and archive date are September 7, 2026; the release body carries a September 8 dateline. Those are two different dates on the same source, not two separate graduation events. Neither date tells you whether a particular installation has passed a recovery exercise.
Check the installed version before borrowing a recipe
Karmada's cluster-failover documentation describes cluster taints, workload eviction, eligible destinations, and scheduling policies. The documentation retrieved for this article is labeled v1.18 and includes a version-specific Failover feature-gate setting. The graduation announcement discusses v1.19. Don't assume a default from one version documents the other.
Use the official documentation that matches your installation. Record the version and relevant configuration with your review evidence, rather than leaving the next operator a link to whatever the documentation website happens to show later.
Define recovery in the customer's terms
Choose one important application action: accepting an order, saving a document, or completing a scheduled job. Write down the observable result that means the action worked. A successful pod placement may be one prerequisite; it isn't the whole acceptance test you've just defined.
For a useful review, ask the application owner to identify what that action depends on. Include the location of its state, the identities it needs, and the systems it must reach. Keep the discussion tied to the selected action. You don't need a diagram of the entire company before you can expose an unanswered recovery question.
Ask these questions before the rehearsal
The following checklist is our proposed review method. It's not a Karmada certification, a claim about a customer's environment, or a guarantee that automation will recover an application.
- Destination: which surviving cluster is eligible, and who has checked that the proposed recovery resources are actually available?
- Dependencies: can the selected workload reach its required network services and obtain the right identity after placement changes?
- State: what is the accepted source of truth, and who decides whether restored or replicated data is suitable for use?
- Duplicate work: what prevents two instances from accepting conflicting writes or repeating an irreversible job during an uncertain transition?
- Decision rights: which steps are automatic, which require approval, and who can stop the exercise or begin rollback?
- Evidence: what will demonstrate that the customer action succeeded, and where will that result be recorded?
Rehearse a bounded failure, then review the gaps
Start with an agreed scenario and an authorized test environment. Set the scope, stopping conditions, observers, and rollback owner before the exercise. Don't interrupt production merely to make a resilience story more convincing. If a safe rehearsal isn't available yet, use a tabletop review and label it as such.
Record what was observed, what was assumed, and what wasn't tested in separate columns. Capture the application result as well as the platform events. When the exercise produces a workaround, document who performed it and whether that person must be available during a real incident. Otherwise the apparent automation may still depend on an undocumented human step.
Buy the review that matches the unanswered question
If your starting problem is understanding restart behavior, health evidence, and recovery ownership for a small service environment, Valen Systems' Reliability Baseline Review is a bounded first step. It reviews supplied evidence for one host and up to three services; it isn't a managed multi-cluster rollout or proof of application failover.
For a larger orchestration environment, talk through the scope before paying for a review that doesn't fit. The goal is a clear next decision: what can already be supported by evidence, what needs a controlled exercise, and who owns the remaining work.
Sources
- CNCF Announces Karmada Graduation — September 7 URL, September 8 datelineSource published September 7, 2026. Retrieved September 9, 2026.
- Karmada: Cluster Failover — v1.18 documentation retrievedPublication date not stated. Retrieved September 9, 2026.
Put this to work.
Talk through the scope with Valen Systems.
A smaller starting point: Review one host first.