Migrating Mutating Webhooks to Admission Policies: A Kubernetes 1.36 Playbook
Kubernetes 1.36 made MutatingAdmissionPolicy stable, so some of your mutating webhooks can be deleted outright. Here is the rule set for deciding which ones go, with the policy YAML and a cutover that does not risk admission.
By VVV Ops ·
Most clusters we audit carry a mutating webhook nobody wants to own. It stamps a label, sets a default resource request, injects an environment variable, and the person who wrote it left two years ago. When its pod goes unready, pod creation hangs for the full ten second timeout and then either fails outright or quietly skips the mutation, depending on a failurePolicy field nobody has read since it was written. Kubernetes 1.36 makes migrating mutating webhooks to admission policies a real option for the first time. Some of those webhooks can go away entirely. Several cannot, and the difference is knowable before you write a line of CEL.
What actually changed in 1.36
Kubernetes v1.36 shipped on 22 April 2026 with 70 enhancements, 18 of them graduating to stable. MutatingAdmissionPolicy is one of the 18. It has been in the tree since v1.30, and as of 1.36 it is stable and enabled by default under admissionregistration.k8s.io/v1.
Three objects make up a policy. MutatingAdmissionPolicy holds the logic. MutatingAdmissionPolicyBinding activates it and scopes where it applies. An optional param resource, any Kubernetes object or CRD referenced through spec.paramKind, supplies values at evaluation time. A policy with no binding is inert, which is a useful property during a migration.
The CEL runs inside the API server process. That is the whole point. A webhook you delete takes a lot with it: a Deployment, a Service, a TLS serving certificate and whatever rotates it, a CA bundle to keep in sync inside the MutatingWebhookConfiguration, a container image to patch every time its base layer gets a CVE, and a pod that has to be running and healthy before any matching object can be created anywhere in the cluster. That last one is why these things page people at 3am.
This is the second 1.36 change worth planning around. The other is the kube-proxy IPVS backend removal, which we covered in the IPVS to nftables migration playbook. Do the network dataplane work first. Admission control can wait a sprint; a broken Service dataplane cannot.
Inventory before you plan anything
Start from what is actually registered, not from what the team remembers deploying:
kubectl get mutatingwebhookconfigurations -o custom-columns=\
"NAME:.metadata.name,\
WEBHOOK:.webhooks[*].name,\
FAIL:.webhooks[*].failurePolicy,\
TIMEOUT:.webhooks[*].timeoutSeconds"
A blank timeout column means the default, and the default for a webhook call is 10 seconds. Any row showing Fail with no explicit timeout is a cluster-wide outage waiting for a bad node.
Sort the results into two piles. Vendor-owned configurations arrived with a Helm chart and will be replaced by the vendor on the vendor's schedule. In-house configurations are yours. In the clusters we look at, teams typically run a handful of mutating webhook configurations and only one or two are their own code. Those one or two are where the entire return on this migration sits.
The four disqualifiers
Work through the in-house pile against this list. One match and the webhook stays.
| Disqualifier | Why the policy cannot do it | Typical example | |---|---|---| | Needs state the API server does not hand you | CEL sees object, oldObject, request, params, namespaceObject, variables and authorizer, and nothing else. No cluster reads, no HTTP calls, no Secret lookups | An injector that pulls credential material from an external secrets store | | Writes an atomic struct, map or array | ApplyConfiguration cannot modify atomic fields, because a merge would risk deleting entries | A webhook that rewrites a whole securityContext block wholesale | | Depends on object metadata outside the accessible set | The reference lists apiVersion, kind, metadata.name, metadata.generateName and metadata.labels as the accessible metadata | A webhook that reads an existing annotation to decide what to do | | Runs imperative logic | CEL has no loops with early exit, no template rendering, no calling into a language runtime | Istio's sidecar injector, which renders a template held in the istio-sidecar-injector ConfigMap |
The third row catches more teams than they expect. Annotations are the most common thing an in-house mutating webhook touches. A policy cannot read one, because the CEL object variable does not expose them, so any webhook that branches on an annotation stays. Writing one is a different matter: a JSONPatch mutation addresses the field by path rather than through the typed ApplyConfiguration object, so a one-way annotation stamp can still move.
A defaulting webhook rewritten as a policy
Here is the shape of the most migratable case: a webhook that sets a default priorityClassName on pods in batch namespaces so that nothing lands in the global default class. In webhook form that is a Deployment, a Service, a cert, and a few dozen lines of Go wrapped around an admission review. In policy form it is this:
# Illustrative policy for an example org (acme.example). Field names and
# patch types are from the MutatingAdmissionPolicy reference; values are ours.
apiVersion: admissionregistration.k8s.io/v1
kind: MutatingAdmissionPolicy
metadata:
name: "default-priority-class.acme.example"
spec:
matchConstraints:
resourceRules:
- apiGroups: [""]
apiVersions: ["v1"]
operations: ["CREATE"]
resources: ["pods"]
matchConditions:
- name: priority-class-not-already-set
expression: "!has(object.spec.priorityClassName) || object.spec.priorityClassName == ''"
failurePolicy: Fail
reinvocationPolicy: IfNeeded
mutations:
- patchType: "ApplyConfiguration"
applyConfiguration:
expression: >
Object{
spec: Object.spec{
priorityClassName: "acme-batch-default"
}
}
---
apiVersion: admissionregistration.k8s.io/v1
kind: MutatingAdmissionPolicyBinding
metadata:
name: "default-priority-class-binding.acme.example"
spec:
policyName: "default-priority-class.acme.example"
matchResources:
namespaceSelector:
matchLabels:
acme.example/tier: batch
Two details in there are load-bearing. The matchConditions guard makes the policy a no-op when the field is already set, which is what lets you run it beside the old webhook during cutover. And reinvocationPolicy: IfNeeded means the policy can be evaluated again after another mutation in the chain changes the object, so the guard has to be genuinely idempotent rather than merely correct on first pass.
One naming rule: policy names ending in .static.k8s.io are reserved for manifest-based admission control, so pick a suffix you control.
Move, keep, or delete
| What it does | Verdict | Reason | |---|---|---| | Defaults priorityClassName, resource requests, imagePullPolicy | Move | Pure function of the incoming object plus params | | Stamps cost-center or team labels from namespace labels | Move | namespaceObject is in the CEL variable set | | Stamps annotations from namespace labels | Move, with JSONPatch | Written by path; the policy cannot read annotations back, so the logic has to stay one-way | | Applies per-tenant tolerations or node selectors from a CRD | Move | This is exactly what paramKind plus a binding is for | | Injects a service mesh sidecar | Keep | Template rendering, vendor-owned, will change under you | | Injects secret material from an external store | Keep | Requires a network call the API server will not make for you | | Rewrites whole securityContext blocks | Keep, or narrow it | Atomic field, and worth splitting into individual field sets first | | Anything whose handler calls a Kubernetes client | Keep | External state, by definition |
We would migrate the top four and leave the rest alone this year. The gain is concentrated: those four are the ones that page you, and they are also the ones with no vendor behind them to fix the pod when it breaks.
Rolling it out without an admission outage
A mutation has no audit mode. It either applies or it does not, so the safety has to come from scoping and from checking the result before you widen.
- Bind to one namespace. Set
matchResources.namespaceSelectoron the binding to a label you put on a single test namespace. The policy is live and the blast radius is one namespace. - Run policy and webhook together. With the
matchConditionsguard above, whichever runs first wins and the other does nothing. There is no flag day. - Read back what the API server would store:
kubectl apply -f pod.yaml --dry-run=server -o yaml. The request goes through the API server without being persisted, so you can diff the returned object against what the webhook used to produce. Do this for the awkward cases, not the happy path: a pod that already sets the field, a pod in a namespace missing the selector label, a pod created by a controller rather than bykubectl. - Widen the namespace selector one tier at a time, then delete the
MutatingWebhookConfigurationwhile leaving the webhook Deployment running for a few days. Rollback at that stage is re-applying a single object. - Delete the Deployment, the Service, the cert, and the cert automation. This is the step teams forget, and it is the step where the maintenance saving actually lands.
Set failurePolicy: Fail on the policy. On a webhook that setting is a loaded gun, because the failure it guards against is a pod being down. On a policy there is no pod, so failing closed costs you nothing and gets you the enforcement guarantee you probably wanted from the webhook all along.
What we would not move yet
Leave vendor webhooks alone. When your service mesh or your certificate manager ships policies, take theirs. Rewriting a vendor's injector in CEL means owning a fork of their admission behaviour across every upgrade, which is a worse job than the one you are trying to get rid of.
Manifest-based admission control is alpha in 1.36. It reads policy files from a staticManifestsDir set in the AdmissionConfiguration passed to --admission-control-config-file, loads them at API server startup before any request is served, and those policies cannot be deleted through the API. It is a good answer to the bootstrap gap and to privileged users removing controls, and it is not something to build a compliance story on while it is alpha. It also explains the reserved .static.k8s.io suffix.
Check your fleet's version floor. The policy API is enabled by default from 1.36, and 1.36 is already available on Amazon EKS, but if one cluster in the fleet is still on 1.34 then the webhook stays until that cluster moves. Running a policy in some clusters and a webhook in others is defensible for a few weeks and miserable for a few months.
Two other things worth doing while you are in here. Confirm every remaining webhook has an explicit timeoutSeconds well below the 10 second default, and confirm none of them intercept the resources their own pods need to start. If your admission stack is the thing standing between a cold cluster and a running control plane, fix that before you optimise anything else. Our Kubernetes production readiness checklist covers the rest of that pre-flight, and the policy-as-code section of our Terraform security guide covers the same enforcement question one layer up, before the manifest ever reaches the cluster.
When to Get Help
This migration is small and self-contained when you have one or two in-house webhooks and a clear picture of what they do. It stops being small when nobody can say what a webhook mutates, when the handler has grown branches for six teams, or when the cluster has enough tenants that a wrong default lands in production before anyone notices.
We do this work as a fixed-scope engagement: inventory every admission configuration in the fleet, classify each one against the disqualifier list, write and test the policies that can move, and leave your team with the runbook for the ones that stay. If you are planning a 1.36 upgrade and want the admission layer sorted as part of it, get in touch.