Multi-Cluster Kubernetes Failover: When It's Worth a Second Control Plane

Karmada just graduated from the CNCF, and every roadmap has multi-cluster failover on it again. What the defaults do to your RTO, what a second control plane costs, and when we tell teams to skip it.

By VVV Ops ·

Your second region has a cluster. It has the same manifests, the same images, roughly the same node pool. Then the primary goes dark and you discover that nothing moves on its own, because multi-cluster Kubernetes failover is a control plane you have to install, configure, and pay for, not a property that emerges from having two clusters. Karmada graduated from the CNCF on 7 September 2026, which is why the question is back on so many roadmaps. This is the version of the answer we give clients: what the defaults actually do to your recovery time, what the second control plane costs, and the cases where we tell teams not to build it.

If you have not finished hardening the single cluster you already run, start with our Kubernetes production readiness checklist and come back. For the business case for spreading across providers at all, our CTO's guide to multi-cloud architecture covers the strategy layer. This post is about the mechanics underneath it.

What graduation changes, and what it does not

Graduation is a governance and maturity signal. The CNCF announcement puts Karmada at more than 1,214 contributors across 292 contributing organizations and more than 5,600 GitHub stars, with a path that ran from a first commit in November 2020 to Sandbox in September 2021 and Incubating in December 2023. Named adopters include Bloomberg, Alibaba Cloud, Huawei, Trip.com, Bilibili, and DaoCloud, among others. The most recent release, v1.19.0, shipped on 31 August 2026.

What graduation does not change is the shape of the work. You still run a second Kubernetes-style control plane with its own API server, its own etcd, its own RBAC surface, and its own upgrade cadence. Member clusters register to it. Workloads are propagated to them by policy rather than applied to them directly. Every operational habit your team has built around one kubeconfig now needs a second answer.

Teams that treat graduation as permission to skip the evaluation are the ones who end up with a federation control plane nobody is on call for.

The three things people mean by failover

Most multi-cluster conversations collapse three separate problems into one word. Separating them decides your architecture.

The first is cluster loss. A control plane or a whole region becomes unreachable and the workloads on it need to run somewhere else. This is what Karmada's cluster failover addresses.

The second is application-level failure. The cluster is healthy, the Deployment is scheduled, but pods will not become ready because a dependency in that region is broken. Karmada handles this separately, under spec.failover.application, because the cluster health signal never fires.

The third is traffic steering. Even after replicas exist in the surviving cluster, something has to send users there. That is DNS, global load balancing, or a service mesh, and Karmada does not do it for you. We have watched a team congratulate itself on a successful failover drill while 100% of production traffic sat on a Route 53 record pointing at the dead region.

Decide which of the three you are buying before you compare tools.

| Capability | Karmada | Argo CD ApplicationSet | DNS or global load balancer | |---|---|---|---| | Deploy the same app to N clusters | Yes | Yes | No | | Split replicas across clusters by weight | Yes | No | No | | Evict and reschedule on cluster health | Yes, when enabled | No | No | | Reschedule when pods never become ready | Yes | No | No | | Move user traffic | No | No | Yes |

The ApplicationSet cluster generator templates one Application per registered cluster. That is a real answer for fan-out deployment and a lot of teams need nothing more. It does not watch cluster health or move replicas, so if your requirement is recovery rather than distribution, it is the wrong tool and adding it will feel like progress for about a quarter.

The defaults put your RTO past six minutes

This is the part missing from every tutorial. Karmada's detection and eviction timings are conservative by default, and if you write an RTO into a customer contract without adding them up, you will miss it.

Working from the karmada-controller-manager reference, health is polled every 5 seconds (--cluster-monitor-period), a running cluster gets 40 seconds of unresponsiveness before it is marked unhealthy (--cluster-monitor-grace-period), and failure has to persist for 30 seconds before the cluster is considered unhealthy (--cluster-failure-threshold). Application failover then waits out decisionConditions.tolerationSeconds, which defaults to 300 seconds.

| Stage | Flag or field | Default | Running total | |---|---|---|---| | Health poll interval | --cluster-monitor-period | 5s | 5s | | Unresponsive grace | --cluster-monitor-grace-period | 40s | 45s | | Failure threshold | --cluster-failure-threshold | 30s | 75s | | Application toleration | tolerationSeconds | 300s | 375s | | Graceful purge window | gracePeriodSeconds | 600s | 975s |

Stack those windows end to end, which is what a bad day does, and you are past six minutes before rescheduling starts. Leave purgeMode on Graciously and the old replicas linger for up to another 600 seconds while the new ones come up. For an active-passive database that overlap is a feature. For anything holding a lease or a singleton lock it is a correctness bug waiting for a bad day.

Then there is throughput. --resource-eviction-rate defaults to 0.5 resources evicted per second in a cluster failover. A member cluster holding 500 propagated resources therefore takes 500 divided by 0.5, or 1,000 seconds, close to 17 minutes, to finish draining. Nothing in the dashboard warns you about this. You find it during the drill, or you find it during the outage.

Our default advice: drop tolerationSeconds to 60 for stateless services, leave it at 300 or higher for anything with a data path, and raise --resource-eviction-rate only after you have measured what the surviving cluster's API server does under that admission load.

Turning cluster failover on is a four-part job

Installing Karmada does not give you failover. The Failover feature gate is Beta and ships as default=false, and the NoExecute eviction path has its own switches. Karmada taints unhealthy clusters with cluster.karmada.io/not-ready when the Ready condition is False and cluster.karmada.io/unreachable when it is Unknown, but those taints carry the NoSchedule effect, which stops new placement and evicts nothing.

To get eviction you enable the gate on the controller manager, allow NoExecute policies on the webhook, and turn on the eviction controller:

--feature-gates=Failover=true              # karmada-controller-manager
--enable-no-execute-taint-eviction=true    # karmada-controller-manager
--allow-no-execute-taint-policy=true       # karmada-webhook

Then the policy itself. This is the application failover shape from the upstream guide, with the toleration tightened for a stateless service:

apiVersion: policy.karmada.io/v1alpha1
kind: PropagationPolicy
metadata:
  name: nginx-propagation
spec:
  failover:
    application:
      decisionConditions:
        tolerationSeconds: 60
      purgeMode: Graciously
      gracePeriodSeconds: 300
  propagateDeps: true
  resourceSelectors:
    - apiVersion: apps/v1
      kind: Deployment
      name: nginx
  placement:
    clusterAffinity:
      clusterNames:
        - member1
        - member2
        - member3
    spreadConstraints:
      - maxGroups: 1
        minGroups: 1
        spreadByField: cluster

propagateDeps: true matters more than it looks. Without it the Deployment lands in the surviving cluster and its ConfigMaps, Secrets, and ServiceAccount do not, and you get a CrashLoopBackOff that reads like an application bug. Test the policy by cordoning a member cluster in a staging fleet, not by reading the YAML and agreeing that it looks right.

What state preservation can and cannot carry

Karmada can carry a small amount of workload state across a failover. statePreservation.rules take a jsonPath into the resource status and an aliasLabelName, and the extracted value is written as a label on the resource in the destination cluster. A Flink job can pick up its checkpoint ID that way. The StatefulFailoverInjection gate that governs it is Alpha and disabled by default.

Read the shape of that carefully. It moves an identifier, not data. Your PersistentVolumes do not follow. If the workload needs its disk, you need storage replication underneath, and that is a database or storage vendor problem with its own RPO, not something the federation layer solves. Any plan that says "Karmada will handle the stateful services" without naming the replication mechanism is not a plan.

For stateful tiers we keep failover out of Karmada entirely. Run the database's own replication and promotion, let Karmada move the stateless services that talk to it, and keep the two decisions separate so a federation bug cannot promote a replica.

What it costs before it saves anything

Control planes are the honest line item. Amazon charges $0.10 per cluster per hour for EKS standard support, or $73 a month at 730 hours. Karmada's own control plane needs a host cluster, so a three-member fleet plus the host is four clusters, which is $292 a month in control planes before a single pod runs. Let a cluster fall onto extended support and that cluster alone is $0.60 per hour, or $438 a month.

The infrastructure cost is the small half. The real cost is that every member cluster is a full cluster: its own upgrades, its own CNI and CSI versions, its own admission controllers, its own certificate rotations. Skew between members is where multi-cluster incidents come from, because the failover works exactly as designed and moves your workload onto a cluster running a CNI two minors behind.

Budget engineering time for a quarterly fleet-wide upgrade cycle and a quarterly failover drill. A federation you have not failed over on purpose in three months is a federation that will not fail over when you need it. If the cluster bill itself is what is driving this conversation, our Kubernetes cost optimization framework is the cheaper first move.

When we tell teams not to do this

Three signals, and any one of them is enough.

You have a single-region compliance boundary. Data residency rules that forbid the workload from running in the second region make the failover target illegal, and no amount of policy YAML fixes that.

Your outages are not cluster outages. Pull the last twelve months of incidents and count how many would have been solved by moving replicas to another cluster. For most teams we work with the honest number is zero, because the failures were bad deploys, expired certificates, exhausted connection pools, and a dependency that was down in every region at once. Multi-cluster failover is expensive insurance against a risk you may not be carrying.

You do not have a traffic story. If DNS failover or a global load balancer is not already in place and tested, the workload arriving in the second cluster is a workload nobody can reach. Solve the routing first. It costs less than a federation and you need it either way.

The teams that get real value here look similar to each other. They run regulated or latency-sensitive workloads across regions they were already paying for, they have hit cluster-level failures more than once, and they have a platform team that can own a second control plane without dropping something else. If that is not you, two independent clusters with GitOps applying the same manifests and DNS in front will get you most of the recovery for a fraction of the operational surface.

When to Get Help

Multi-cluster Kubernetes failover fails in the gap between the diagram and the defaults. The architecture review is easy. Adding up detection windows, eviction rates, and purge modes against a contractual RTO, then proving it in a drill, is the part that takes a team that has done it before.

We help engineering teams decide whether federation is warranted, size the control plane and operational cost honestly, and run the first failover drill against a staging fleet before anything touches production. If you are staring at a Karmada evaluation and want a second opinion on whether it is the right answer, talk to us.

Tags: multi-cluster kubernetes failover, karmada multi-cluster orchestration, kubernetes deployment automation, multi-cloud architecture best practices, cloud-native architecture, secure kubernetes deployments