Azure Kubernetes Service Backup: Complete Production Guide
Abdullah
2026-08-09
Kubernetes & DevOps
<div class="blog-content">
<section class="intro">
<h2>Why Your AKS Clusters Need Backup Today</h2>
<p>You've built a solid AKS cluster. Workloads are humming. CI/CD pipelines running smoothly. <strong>Then disaster strikes.</strong></p>
<p>A developer accidentally deletes a critical namespace. A corrupted CRD cascades failures. Storage corrupts silently. A bad upgrade breaks everything.</p>
<div class="alert alert-warning">
<div class="alert-icon">⚠️</div>
<div class="alert-content">
<strong>Critical Reality Check:</strong> Without backups, you're not recovering—you're rebuilding from scratch. I've watched teams lose entire days to infrastructure recreation that could've been restored in hours.
</div>
</div>
<p>AKS backup isn't optional. <strong>It's your disaster insurance.</strong></p>
</section>
<section class="challenges">
<h2>The Hard Problems (Why This Matters)</h2>
<div class="challenge-box">
<div class="box-header">
<span class="box-number">1</span>
<h3>Your Cluster State is Fragmented</h3>
</div>
<p>Kubernetes doesn't store everything in one place. You have:</p>
<ul>
<li><strong>Cluster Config:</strong> Node pools, storage classes, network policies</li>
<li><strong>Application State:</strong> Deployments, StatefulSets, Secrets, CRDs</li>
<li><strong>Persistent Data:</strong> Azure Disks, Azure Files, database volumes</li>
</ul>
<div class="alert alert-info">
<strong>Important:</strong> Each layer needs different backup strategies. You can't just backup etcd and pray.
</div>
</div>
<div class="challenge-box">
<div class="box-header">
<span class="box-number">2</span>
<h3>Workload Identity Breaks After Restore</h3>
</div>
<p>Modern AKS uses Workload Identity federation. Your pods authenticate to Azure without credentials. <strong>But here's the trap:</strong></p>
<ul>
<li>OIDC issuer URLs change on every cluster</li>
<li>Federated credentials trust specific URLs only</li>
<li>Restore to new cluster = different URL = broken authentication</li>
</ul>
<div class="alert alert-danger">
<strong>Silent Failure Risk:</strong> Your apps restore fine but can't talk to Azure services. Silent failures everywhere.
</div>
</div>
<div class="challenge-box">
<div class="box-header">
<span class="box-number">3</span>
<h3>Volume Snapshots Require Setup Most Teams Skip</h3>
</div>
<p>Azure Data Protection needs CSI snapshot drivers. <strong>Most AKS clusters don't have them installed.</strong></p>
<ul>
<li>AKS doesn't enable volume snapshots by default</li>
<li>Backup policies fail silently when drivers are missing</li>
<li>You only discover this during an actual disaster</li>
</ul>
<div class="callout callout-attention">
<p><strong>Detection Problem:</strong> Silent failures in backup systems are the worst kind—discovered only in crisis.</p>
</div>
</div>
<div class="challenge-box">
<div class="box-header">
<span class="box-number">4</span>
<h3>Testing Backups is Expensive and Hard</h3>
</div>
<p>The worst time to discover your backups are corrupted is during a real incident. <strong>But testing means:</strong></p>
<ul>
<li>Spinning up temporary restore clusters</li>
<li>Handling Workload Identity re-configuration</li>
<li>Validating every app can authenticate</li>
<li>Cleaning up resources to save costs</li>
</ul>
<div class="alert alert-warning">
<strong>Team Reality:</strong> Most teams skip this entirely. Then catastrophe happens.
</div>
</div>
</section>
<section class="incidents">
<h2>5 Production Incidents We Faced (And How We Fixed Them)</h2>
<div class="incident">
<div class="incident-header">
<span class="incident-badge">INCIDENT #1</span>
<h3>"Upgrade Override Cannot Be Unset" Error</h3>
</div>
<p><strong>What happened:</strong> Terraform refused to remove upgrade_override settings</p>
<p>Azure tracks configuration state with strict rules. Once you enable auto-upgrade, removing it requires special handling.</p>
<div class="code-block">
<div class="code-header">Terraform Fix</div>
<pre><code>lifecycle {
ignore_changes = [
kubernetes_cluster_config.0.upgrade_override
]
}</code></pre>
</div>
<div class="lesson-box">
<span class="lesson-icon">💡</span>
<strong>Lesson:</strong> AKS configuration has cascading effects. Test everything in dev first.
</div>
</div>
<div class="incident">
<div class="incident-header">
<span class="incident-badge">INCIDENT #2</span>
<h3>PodDisruptionBudget Blocked Istio Upgrade for Hours</h3>
</div>
<p><strong>The nightmare:</strong> Upgrading Istio from v1.28 to v1.29 hung indefinitely. PodDisruptionBudgets were too restrictive—the system couldn't evict pods.</p>
<div class="alert alert-danger">
<strong>Production Impact:</strong> 3+ hours of downtime. Complete deployment blockage.
</div>
<div class="code-block">
<div class="code-header">Emergency Fix (kubectl)</div>
<pre><code>kubectl get pdb -A -o yaml > pdb-backup.yaml
kubectl delete pdb -n istio-system istio-ingressgateway
# Upgrade completes
kubectl apply -f pdb-backup.yaml</code></pre>
</div>
<div class="code-block">
<div class="code-header">Better Approach (YAML)</div>
<pre><code>apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: istio-ingressgateway
spec:
minAvailable: 1 # Allow some unavailability during maintenance
selector:
matchLabels:
app: istio-ingressgateway</code></pre>
</div>
<div class="lesson-box">
<span class="lesson-icon">💡</span>
<strong>Lesson:</strong> PDBs are safety mechanisms, but overly restrictive ones become roadblocks. Always allow some unavailability.
</div>
</div>
<div class="incident">
<div class="incident-header">
<span class="incident-badge">INCIDENT #3</span>
<h3>Volume Snapshots Silently Failed (CSI Driver Missing)</h3>
</div>
<p><strong>The problem:</strong> Backups appeared to succeed but snapshots never created. AKS doesn't install volume snapshot CSI drivers by default.</p>
<div class="alert alert-warning">
<strong>Silent Failure:</strong> No errors in logs. No failed jobs. Just... no snapshots.
</div>
<div class="code-block">
<div class="code-header">Fix via Terraform</div>
<pre><code>resource "azurerm_kubernetes_cluster" "main" {
storage_profile {
snapshot_controller_enabled = true # ← This was missing!
}
}</code></pre>
</div>
<div class="lesson-box">
<span class="lesson-icon">💡</span>
<strong>Lesson:</strong> Storage infrastructure is invisible until it breaks. Verify before configuring backups.
</div>
</div>
<div class="incident">
<div class="incident-header">
<span class="incident-badge">INCIDENT #4</span>
<h3>OIDC Issuer Mismatch Broke All Pod Authentication</h3>
</div>
<p><strong>The horror:</strong> Restored cluster worked. Apps started. Then everything failed silently.</p>
<div class="alert alert-danger">
<strong>Error in Logs:</strong> <code>AADSTS700211: No matching federated identity record</code>
</div>
<p><strong>Why:</strong> Each cluster gets a unique OIDC issuer URL. Federated credentials trust specific URLs only. Restore = new URL = broken trust.</p>
<div class="code-block">
<div class="code-header">Fix: Update Federated Credentials</div>
<pre><code># Get new cluster's OIDC issuer
NEW_OIDC=$(az aks show -g my-rg -n my-cluster \
--query "oidcIssuerProfile.issuerUrl" -otsv)
# Update federated credential to trust new issuer
az ad app federated-credential create \
--id <AppObjectId> \
--parameters "{
'issuer': '$NEW_OIDC',
'subject': 'system:serviceaccount:default:default',
'audiences': ['api://AzureADTokenExchange']
}"</code></pre>
</div>
<div class="lesson-box">
<span class="lesson-icon">💡</span>
<strong>Lesson:</strong> Workload Identity is powerful but fragile. After any restore, audit all OIDC-dependent workloads immediately.
</div>
</div>
<div class="incident">
<div class="incident-header">
<span class="incident-badge">INCIDENT #5</span>
<h3>Permissions Failed During Restore</h3>
</div>
<p><strong>The issue:</strong> Backup succeeded. Restore failed with permission denied.</p>
<p>Your AKS cluster's managed identity needs explicit permissions on the snapshot resource group. <strong>Permissions aren't inherited.</strong></p>
<div class="code-block">
<div class="code-header">Fix: Add RBAC Role Assignment</div>
<pre><code>resource "azurerm_role_assignment" "aks_snapshot" {
scope = azurerm_resource_group.snapshots.id
role_definition_name = "Contributor"
principal_id = azurerm_kubernetes_cluster.main.kubelet_identity[0].object_id
}</code></pre>
</div>
<div class="lesson-box">
<span class="lesson-icon">💡</span>
<strong>Lesson:</strong> Always explicitly grant RBAC roles. Never assume inheritance.
</div>
</div>
</section>
<section class="terraform-section">
<h2>Production-Ready Terraform Module</h2>
<div class="callout callout-tip">
<strong>💼 Production Grade:</strong> This is copy-paste ready for your infrastructure. Use this. Don't reinvent it.
</div>
<div class="code-block">
<div class="code-header">modules/aks_backup/main.tf</div>
<pre><code class="language-hcl">terraform {
required_providers {
azurerm = {
source = "hashicorp/azurerm"
version = "~> 3.80"
}
}
}
variable "cluster_name" { type = string }
variable "cluster_id" { type = string }
variable "location" { type = string }
variable "resource_group_name" { type = string }
# ===== BACKUP VAULT =====
resource "azurerm_data_protection_backup_vault" "main" {
name = "${var.cluster_name}-vault"
resource_group_name = var.resource_group_name
location = var.location
datastore_type = "VaultStore"
redundancy = "GeoRedundant"
identity {
type = "SystemAssigned"
}
tags = {
Environment = "production"
Purpose = "aks-cluster-backup"
}
}
# ===== BACKUP POLICY =====
resource "azurerm_data_protection_backup_policy_kubernetes_cluster" "main" {
name = "${var.cluster_name}-backup-policy"
resource_group_name = var.resource_group_name
vault_name = azurerm_data_protection_backup_vault.main.name
backup_repeating_time_intervals = ["R/2026-01-01T00:00:00Z/PT4H"]
default_retention_rule {
lifetime = "P7D"
}
retention_rule {
name = "weekly-retention"
lifetime = "P4W"
criteria {
days_of_week = ["Sunday"]
scheduled_backup_times = ["2026-01-01T02:00:00Z"]
}
}
}
# ===== TRUSTED ACCESS =====
resource "azurerm_role_assignment" "vault_aks_backup" {
scope = var.cluster_id
role_definition_name = "Kubernetes Service Cluster User Role"
principal_id = azurerm_data_protection_backup_vault.main.identity[0].principal_id
}
# ===== OUTPUTS =====
output "backup_vault_id" {
value = azurerm_data_protection_backup_vault.main.id
}</code></pre>
</div>
<div class="code-block">
<div class="code-header">Usage: main.tf</div>
<pre><code>module "aks_backup" {
source = "./modules/aks_backup"
cluster_name = azurerm_kubernetes_cluster.prod.name
cluster_id = azurerm_kubernetes_cluster.prod.id
location = azurerm_resource_group.main.location
resource_group_name = azurerm_resource_group.main.name
}</code></pre>
</div>
</section>
<section class="decisions">
<h2>3 Critical Configuration Decisions</h2>
<div class="decision-card">
<h3>Decision 1: When to Backup (Frequency Matters)</h3>
<div class="comparison-table">
<div class="comparison-row">
<div class="comparison-item">
<strong>Every 4 Hours</strong>
<p class="text-muted">Higher cost, lower RPO (4 hours max data loss)</p>
</div>
<div class="comparison-item">
<strong>Daily</strong>
<p class="text-muted">Lower cost, higher RPO (24 hours max data loss)</p>
</div>
</div>
</div>
<div class="alert alert-info">
<strong>Right Answer:</strong> Match your business needs. Don't over-engineer costs.
</div>
</div>
<div class="decision-card">
<h3>Decision 2: Redundancy Strategy (Safety vs Cost)</h3>
<div class="comparison-table">
<div class="comparison-row">
<div class="comparison-item">
<strong>Production</strong>
<p class="text-muted">GeoRedundant (survives regional Azure outages)</p>
</div>
<div class="comparison-item">
<strong>Development</strong>
<p class="text-muted">LocallyRedundant (saves 30% on backups)</p>
</div>
</div>
</div>
</div>
<div class="decision-card">
<h3>Decision 3: What NOT to Backup</h3>
<div class="callout callout-warning">
<strong>Don't backup kube-system namespace.</strong> System components are ephemeral and reproducible.
</div>
<p><strong>Why:</strong> Saves 20-30% of snapshot costs. Backup only application data.</p>
</div>
</section>
<section class="costs">
<h2>What Does It Actually Cost?</h2>
<div class="pricing-table">
<table>
<thead>
<tr>
<th>Component</th>
<th>Monthly Cost</th>
<th>Notes</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>Backup Vault</strong></td>
<td class="cost-primary">$50-100</td>
<td>Flat rate per vault</td>
</tr>
<tr>
<td><strong>Snapshots</strong></td>
<td class="cost-primary">$0.05-0.15/GB</td>
<td>Incremental + deduplication</td>
</tr>
<tr>
<td><strong>Storage Logs</strong></td>
<td class="cost-primary">$0.02-0.05/GB</td>
<td>Minimal overhead</td>
</tr>
<tr class="total-row">
<td><strong>TOTAL</strong></td>
<td class="cost-total"><strong>$200-500</strong></td>
<td>Varies by cluster size</td>
</tr>
</tbody>
</table>
</div>
<div class="cost-savings-box">
<h4>💰 How to Save Money</h4>
<ul>
<li><strong>Exclude system namespaces</strong> → 30% savings</li>
<li><strong>Daily backups for dev</strong> → Lower costs</li>
<li><strong>Shorter retention for dev</strong> → 7 days vs 30 days</li>
<li><strong>LocallyRedundant for dev</strong> → 30% cost reduction</li>
</ul>
</div>
</section>
<section class="mistakes">
<h2>5 Mistakes That Will Hurt You</h2>
<div class="mistake-box">
<div class="mistake-header">
<span class="mistake-icon">❌</span>
<h3>Mistake 1: Never Testing Restores</h3>
</div>
<div class="mistake-content">
<p><strong>Impact:</strong> Corrupted backups discovered during actual disaster</p>
<div class="alert alert-danger">Worst case: Complete data loss discovered only when needed</div>
<p><strong>Fix:</strong> Do monthly restore tests to temporary cluster</p>
</div>
</div>
<div class="mistake-box">
<div class="mistake-header">
<span class="mistake-icon">❌</span>
<h3>Mistake 2: Backing Up Everything</h3>
</div>
<div class="mistake-content">
<p><strong>Impact:</strong> Snapshot costs spiral. Backups slow. Unnecessary data.</p>
<p><strong>Fix:</strong> Exclude system namespaces explicitly</p>
</div>
</div>
<div class="mistake-box">
<div class="mistake-header">
<span class="mistake-icon">❌</span>
<h3>Mistake 3: Forgetting Workload Identity After Restore</h3>
</div>
<div class="mistake-content">
<p><strong>Impact:</strong> Apps can't authenticate. Silent failures everywhere.</p>
<p><strong>Fix:</strong> Document OIDC issuer update in your runbook</p>
</div>
</div>
<div class="mistake-box">
<div class="mistake-header">
<span class="mistake-icon">❌</span>
<h3>Mistake 4: No Monitoring of Backup Jobs</h3>
</div>
<div class="mistake-content">
<p><strong>Impact:</strong> Backups fail silently for weeks</p>
<p><strong>Fix:</strong> Set up Azure Monitor alerts for failed jobs</p>
</div>
</div>
<div class="mistake-box">
<div class="mistake-header">
<span class="mistake-icon">❌</span>
<h3>Mistake 5: One Vault for All Regions</h3>
</div>
<div class="mistake-content">
<p><strong>Impact:</strong> Regional Azure outage breaks both production and backups</p>
<p><strong>Fix:</strong> One vault per cluster, use geo-redundancy</p>
</div>
</div>
</section>
<section class="runbook">
<h2>Disaster Recovery Runbook (Namespace Deletion)</h2>
<div class="callout callout-important">
<strong>⏱️ Total Recovery Time: ~30 minutes</strong>
</div>
<div class="runbook-step">
<div class="step-header">
<span class="step-number">1</span>
<h3>Assess (5 minutes)</h3>
</div>
<div class="code-block">
<pre><code>kubectl get ns | grep deleted-namespace
az dataprotection backup-instance show \
--vault-name my-vault \
--resource-group my-rg \
--backup-instance-name my-backup</code></pre>
</div>
</div>
<div class="runbook-step">
<div class="step-header">
<span class="step-number">2</span>
<h3>Restore to Temporary Cluster (15 minutes)</h3>
</div>
<div class="code-block">
<pre><code>az aks create \
--resource-group my-rg \
--name temp-restore-cluster \
--node-count 2
az dataprotection job create \
--vault-name my-vault \
--resource-group my-rg \
--backup-instance-name my-backup-instance
# Wait 5-10 minutes for restore to complete</code></pre>
</div>
</div>
<div class="runbook-step">
<div class="step-header">
<span class="step-number">3</span>
<h3>Extract and Reapply (10 minutes)</h3>
</div>
<div class="code-block">
<pre><code>kubectl config use-context temp-restore-cluster
kubectl get all -n deleted-namespace -o yaml > namespace-backup.yaml
kubectl config use-context prod-cluster
kubectl apply -f namespace-backup.yaml</code></pre>
</div>
</div>
<div class="runbook-step">
<div class="step-header">
<span class="step-number">4</span>
<h3>Cleanup (2 minutes)</h3>
</div>
<div class="code-block">
<pre><code>az aks delete --name temp-restore-cluster --resource-group my-rg --yes</code></pre>
</div>
</div>
</section>
<section class="conclusion">
<h2>The Bottom Line</h2>
<div class="conclusion-box">
<h3>🛡️ AKS backup isn't optional. It's insurance against catastrophe.</h3>
<p>The teams that invest in proper backup strategy now are the ones who sleep soundly later. The teams that skip it? They're the ones rebuilding infrastructure at 3 AM.</p>
</div>
<h3>5 Things to Do Right Now</h3>
<div class="action-list">
<div class="action-item">
<span class="action-number">1</span>
<p>Define your RPO/RTO (Recovery Point Objective / Recovery Time Objective)</p>
</div>
<div class="action-item">
<span class="action-number">2</span>
<p>Deploy this Terraform module to production TODAY</p>
</div>
<div class="action-item">
<span class="action-number">3</span>
<p>Schedule your first restore test</p>
</div>
<div class="action-item">
<span class="action-number">4</span>
<p>Set up Azure Monitor alerts for failed backups</p>
</div>
<div class="action-item">
<span class="action-number">5</span>
<p>Document your Workload Identity OIDC update procedure</p>
</div>
</div>
<div class="alert alert-success">
<strong>Remember:</strong> The investment pays dividends the moment disaster strikes. Don't be the team rebuilding from scratch.
</div>
</section>
</div>
← Back to Blog