k8s-security.pro
kubernetes security aks azure devops

AKS Security Best Practices: The Checklist I Actually Use

AKS security best practices: workload identity, IMDS exposure, local accounts, private API server, network policy engines, and the audit logs that are off.

K8s Security Pro Team | | 10 min read

AKS Security Best Practices: The Checklist I Actually Use

Of the three big managed Kubernetes services, AKS is the one where the gap between “cluster that works” and “cluster that would pass an audit” surprises people the most. Azure runs the control plane fine. But AKS ships with a public API server, working local admin credentials, no enforced network policy, and audit logging you have to assemble yourself from diagnostic settings.

None of this is exotic. It’s the same short list of items, cluster after cluster. Here it is, ordered by how bad things get when each one is wrong.

1. Disable local accounts

This is the AKS-specific finding I put first, because it undoes everything else you configure. AKS clusters come with certificate-based local admin access, and it keeps working even after you integrate Entra ID. Anyone with the right Azure role can run:

az aks get-credentials --name my-cluster --resource-group my-rg --admin

and walk straight past Entra ID, Conditional Access, MFA, and your Kubernetes RBAC, with a long-lived credential (valid for years) that cannot be individually revoked.

The fix is one flag, once Entra ID integration is in place:

az aks update --name my-cluster --resource-group my-rg --disable-local-accounts

Check where you stand first:

az aks show --name my-cluster --resource-group my-rg \
  --query "{localAccounts: disableLocalAccounts, aad: aadProfile.managed}"

If localAccounts comes back false or empty, that side door is open.

2. Workload identity for pods, and mind the metadata endpoint

Pods that need Azure access should get it through Microsoft Entra Workload ID: the pod’s Kubernetes service account federates to an Entra identity, tokens are short-lived, and no secret is stored anywhere. If you are still using the old pod-identity add-on or, worse, service principal secrets mounted into pods, migrating is the highest-value identity work you can do. And if you find exported service principal credentials in a Secret, treat them like the leak they will eventually become.

The part people miss: AKS nodes have a managed identity, and pods can reach the instance metadata endpoint at 169.254.169.254 to request tokens for it. One SSRF bug in one app and an attacker is holding Azure credentials, not just a shell. Microsoft’s own hardening guidance says to block pod access to IMDS with a network policy, which conveniently requires you to have done item 4 below.

3. Stop exposing the API server

Same story as EKS and GKE, same priority: new AKS clusters answer on a public endpoint. Pick your remedy by how your team works. Authorized IP ranges is the ten-minute fix:

az aks update --name my-cluster --resource-group my-rg \
  --api-server-authorized-ip-ranges 203.0.113.0/24

A private cluster (or API Server VNet integration) is the thorough fix if your engineers are already on a VPN or ExpressRoute. Either way, “authenticated but reachable from every IP on the internet” is the state to leave.

4. Make sure something actually enforces your NetworkPolicies

AKS has the same silent trap as GKE, arguably worse: the network policy engine is a cluster configuration choice, and if none was chosen, every NetworkPolicy you apply is accepted and ignored. No error, no event, nothing. I have seen a cluster with a beautiful set of default-deny policies and no engine to read them.

Check what you are running:

az aks show --name my-cluster --resource-group my-rg \
  --query "networkProfile.networkPolicy"

Empty or none means your policies are decorative. The good news: AKS now lets you enable a network policy engine on existing clusters, and Azure CNI powered by Cilium is the option I default to for new ones. Then the usual drill: default-deny per namespace, explicit allows, and the DNS exception to kube-system before anything else, plus the IMDS block from item 2.

5. Turn on the audit logs (and watch the cost)

AKS control plane logs, including kube-audit, exist but go nowhere until you create a diagnostic setting pointing them at a Log Analytics workspace. Out of the box, “who deleted that deployment” has no answer.

Two practical notes. First, enable kube-audit-admin rather than full kube-audit if cost matters: it captures create, update, and delete while dropping the flood of read events, which is most of the volume. One caveat that matters for security: that also means it does not record who read a Secret. If your threat model or compliance framework needs that answer, only full kube-audit has it. Second, set retention deliberately. The audit trail you need in an incident is usually weeks old by the time you need it.

While you are in that blade, send kube-apiserver and guard (the Entra ID auth component) logs along too. They are what turn “someone authenticated” into “this person, from this IP”.

6. Entra ID for RBAC, Azure RBAC for the cluster boundary

With local accounts closed, make Entra ID do the work: groups mapped to Kubernetes RBAC roles (or Azure RBAC for Kubernetes authorization if you want role assignments managed entirely on the Azure side), engineers in groups, no standing cluster-admin for humans. The audit story gets dramatically better too, since every action ties back to a real identity that offboarding actually removes.

The check that finds the skeletons:

kubectl get clusterrolebindings -o json | jq -r '.items[] | select(.roleRef.name=="cluster-admin") | .metadata.name'

Anything in that list that is not a system binding deserves an explanation.

7. The in-cluster baseline is still yours

Managed control plane, unmanaged habits, third time in this series: nothing in AKS stops privileged containers, hostPath mounts, :latest tags, missing resource limits, or service account tokens automounted into pods that never call the API. Pod Security Admission per namespace, or Kyverno or Gatekeeper if you need policy with exceptions, plus the boring RBAC hygiene above.

For a fast pass over exactly this category I use k8s-audit, an open source script that runs 16 read-only checks with kubectl and jq in about 30 seconds. It will not replace a CIS scan; it tells you which fires are real before you open the 400-line report.

8. Odds and ends that auditors ask about

Defender for Containers if you have compliance requirements or meaningful exposure, it consolidates runtime detection and vulnerability assessment you would otherwise stitch together. An auto-upgrade channel (patch at minimum) so the cluster does not quietly age out of support. Key Vault via the Secrets Store CSI driver for secrets that matter, or better, External Secrets Operator so rotation is real. And KMS etcd encryption with your own key if your compliance framework asks who controls the encryption key, because the platform default answer is “Microsoft”.

Where this fits in an audit

Items 1 through 6 map to CIS AKS Benchmark controls, and the az commands above produce most of the evidence. The in-cluster half is identical across EKS, GKE, and AKS, which is why the same checklist keeps working.

That is what our 50-point checklist packages: every item with the check command, the fix, and the CIS/SOC2 mapping. And if you would rather have it done for you, the first three customers of our audit service get the full review for $99 in exchange for honest feedback. You run the read-only scans, I turn the output into a prioritized fix plan.

FAQ

What should I fix first on an existing AKS cluster?

Local accounts, then the API server exposure. Both are single az commands, and together they close the two paths that turn a leaked credential into a cluster compromise.

Does Entra Workload ID replace pod-managed identity?

Yes. The old aad-pod-identity add-on is deprecated, and Entra Workload ID (OIDC federation) is the supported path. Migration mostly means new service account annotations and an identity federation, no secrets involved.

Do Azure NSGs replace NetworkPolicies?

No. NSGs govern subnet and NIC traffic; pod-to-pod traffic inside the cluster needs NetworkPolicy, enforced by an engine the cluster actually has. Both layers, different jobs.

How do I check all of this quickly?

The az commands in this post are read-only queries. For the in-cluster half, k8s-audit gives you the 30-second pass, then kube-bench with the AKS benchmark for depth.

Want this handled for you?

Done-for-you cluster audit with a prioritized, CIS-mapped fix plan. First 3 customers: $99 pilot. Or grab the free kit and do it yourself.

Get the $99 Pilot Audit

Get the Free K8s Security Quick-Start Kit

Get 5 essential templates + audit checklist highlights delivered to your inbox.

No spam. Unsubscribe anytime.

Secure Your Kubernetes Clusters

Get the complete 50-point audit checklist and 20+ production-ready YAML templates.

View Pricing Plans