Azure Kubernetes Service Backup: Complete Production Guide

Abdullah
2026-08-09
Kubernetes & DevOps
<div class="blog-content"> <section class="intro"> <h2>Why Your AKS Clusters Need Backup Today</h2> <p>You've built a solid AKS cluster. Workloads are humming. CI/CD pipelines running smoothly. <strong>Then disaster strikes.</strong></p> <p>A developer accidentally deletes a critical namespace. A corrupted CRD cascades failures. Storage corrupts silently. A bad upgrade breaks everything.</p> <div class="alert alert-warning"> <div class="alert-icon">⚠️</div> <div class="alert-content"> <strong>Critical Reality Check:</strong> Without backups, you're not recovering—you're rebuilding from scratch. I've watched teams lose entire days to infrastructure recreation that could've been restored in hours. </div> </div> <p>AKS backup isn't optional. <strong>It's your disaster insurance.</strong></p> </section> <section class="challenges"> <h2>The Hard Problems (Why This Matters)</h2> <div class="challenge-box"> <div class="box-header"> <span class="box-number">1</span> <h3>Your Cluster State is Fragmented</h3> </div> <p>Kubernetes doesn't store everything in one place. You have:</p> <ul> <li><strong>Cluster Config:</strong> Node pools, storage classes, network policies</li> <li><strong>Application State:</strong> Deployments, StatefulSets, Secrets, CRDs</li> <li><strong>Persistent Data:</strong> Azure Disks, Azure Files, database volumes</li> </ul> <div class="alert alert-info"> <strong>Important:</strong> Each layer needs different backup strategies. You can't just backup etcd and pray. </div> </div> <div class="challenge-box"> <div class="box-header"> <span class="box-number">2</span> <h3>Workload Identity Breaks After Restore</h3> </div> <p>Modern AKS uses Workload Identity federation. Your pods authenticate to Azure without credentials. <strong>But here's the trap:</strong></p> <ul> <li>OIDC issuer URLs change on every cluster</li> <li>Federated credentials trust specific URLs only</li> <li>Restore to new cluster = different URL = broken authentication</li> </ul> <div class="alert alert-danger"> <strong>Silent Failure Risk:</strong> Your apps restore fine but can't talk to Azure services. Silent failures everywhere. </div> </div> <div class="challenge-box"> <div class="box-header"> <span class="box-number">3</span> <h3>Volume Snapshots Require Setup Most Teams Skip</h3> </div> <p>Azure Data Protection needs CSI snapshot drivers. <strong>Most AKS clusters don't have them installed.</strong></p> <ul> <li>AKS doesn't enable volume snapshots by default</li> <li>Backup policies fail silently when drivers are missing</li> <li>You only discover this during an actual disaster</li> </ul> <div class="callout callout-attention"> <p><strong>Detection Problem:</strong> Silent failures in backup systems are the worst kind—discovered only in crisis.</p> </div> </div> <div class="challenge-box"> <div class="box-header"> <span class="box-number">4</span> <h3>Testing Backups is Expensive and Hard</h3> </div> <p>The worst time to discover your backups are corrupted is during a real incident. <strong>But testing means:</strong></p> <ul> <li>Spinning up temporary restore clusters</li> <li>Handling Workload Identity re-configuration</li> <li>Validating every app can authenticate</li> <li>Cleaning up resources to save costs</li> </ul> <div class="alert alert-warning"> <strong>Team Reality:</strong> Most teams skip this entirely. Then catastrophe happens. </div> </div> </section> <section class="incidents"> <h2>5 Production Incidents We Faced (And How We Fixed Them)</h2> <div class="incident"> <div class="incident-header"> <span class="incident-badge">INCIDENT #1</span> <h3>"Upgrade Override Cannot Be Unset" Error</h3> </div> <p><strong>What happened:</strong> Terraform refused to remove upgrade_override settings</p> <p>Azure tracks configuration state with strict rules. Once you enable auto-upgrade, removing it requires special handling.</p> <div class="code-block"> <div class="code-header">Terraform Fix</div> <pre><code>lifecycle { ignore_changes = [ kubernetes_cluster_config.0.upgrade_override ] }</code></pre> </div> <div class="lesson-box"> <span class="lesson-icon">💡</span> <strong>Lesson:</strong> AKS configuration has cascading effects. Test everything in dev first. </div> </div> <div class="incident"> <div class="incident-header"> <span class="incident-badge">INCIDENT #2</span> <h3>PodDisruptionBudget Blocked Istio Upgrade for Hours</h3> </div> <p><strong>The nightmare:</strong> Upgrading Istio from v1.28 to v1.29 hung indefinitely. PodDisruptionBudgets were too restrictive—the system couldn't evict pods.</p> <div class="alert alert-danger"> <strong>Production Impact:</strong> 3+ hours of downtime. Complete deployment blockage. </div> <div class="code-block"> <div class="code-header">Emergency Fix (kubectl)</div> <pre><code>kubectl get pdb -A -o yaml > pdb-backup.yaml kubectl delete pdb -n istio-system istio-ingressgateway # Upgrade completes kubectl apply -f pdb-backup.yaml</code></pre> </div> <div class="code-block"> <div class="code-header">Better Approach (YAML)</div> <pre><code>apiVersion: policy/v1 kind: PodDisruptionBudget metadata: name: istio-ingressgateway spec: minAvailable: 1 # Allow some unavailability during maintenance selector: matchLabels: app: istio-ingressgateway</code></pre> </div> <div class="lesson-box"> <span class="lesson-icon">💡</span> <strong>Lesson:</strong> PDBs are safety mechanisms, but overly restrictive ones become roadblocks. Always allow some unavailability. </div> </div> <div class="incident"> <div class="incident-header"> <span class="incident-badge">INCIDENT #3</span> <h3>Volume Snapshots Silently Failed (CSI Driver Missing)</h3> </div> <p><strong>The problem:</strong> Backups appeared to succeed but snapshots never created. AKS doesn't install volume snapshot CSI drivers by default.</p> <div class="alert alert-warning"> <strong>Silent Failure:</strong> No errors in logs. No failed jobs. Just... no snapshots. </div> <div class="code-block"> <div class="code-header">Fix via Terraform</div> <pre><code>resource "azurerm_kubernetes_cluster" "main" { storage_profile { snapshot_controller_enabled = true # ← This was missing! } }</code></pre> </div> <div class="lesson-box"> <span class="lesson-icon">💡</span> <strong>Lesson:</strong> Storage infrastructure is invisible until it breaks. Verify before configuring backups. </div> </div> <div class="incident"> <div class="incident-header"> <span class="incident-badge">INCIDENT #4</span> <h3>OIDC Issuer Mismatch Broke All Pod Authentication</h3> </div> <p><strong>The horror:</strong> Restored cluster worked. Apps started. Then everything failed silently.</p> <div class="alert alert-danger"> <strong>Error in Logs:</strong> <code>AADSTS700211: No matching federated identity record</code> </div> <p><strong>Why:</strong> Each cluster gets a unique OIDC issuer URL. Federated credentials trust specific URLs only. Restore = new URL = broken trust.</p> <div class="code-block"> <div class="code-header">Fix: Update Federated Credentials</div> <pre><code># Get new cluster's OIDC issuer NEW_OIDC=$(az aks show -g my-rg -n my-cluster \ --query "oidcIssuerProfile.issuerUrl" -otsv) # Update federated credential to trust new issuer az ad app federated-credential create \ --id <AppObjectId> \ --parameters "{ 'issuer': '$NEW_OIDC', 'subject': 'system:serviceaccount:default:default', 'audiences': ['api://AzureADTokenExchange'] }"</code></pre> </div> <div class="lesson-box"> <span class="lesson-icon">💡</span> <strong>Lesson:</strong> Workload Identity is powerful but fragile. After any restore, audit all OIDC-dependent workloads immediately. </div> </div> <div class="incident"> <div class="incident-header"> <span class="incident-badge">INCIDENT #5</span> <h3>Permissions Failed During Restore</h3> </div> <p><strong>The issue:</strong> Backup succeeded. Restore failed with permission denied.</p> <p>Your AKS cluster's managed identity needs explicit permissions on the snapshot resource group. <strong>Permissions aren't inherited.</strong></p> <div class="code-block"> <div class="code-header">Fix: Add RBAC Role Assignment</div> <pre><code>resource "azurerm_role_assignment" "aks_snapshot" { scope = azurerm_resource_group.snapshots.id role_definition_name = "Contributor" principal_id = azurerm_kubernetes_cluster.main.kubelet_identity[0].object_id }</code></pre> </div> <div class="lesson-box"> <span class="lesson-icon">💡</span> <strong>Lesson:</strong> Always explicitly grant RBAC roles. Never assume inheritance. </div> </div> </section> <section class="terraform-section"> <h2>Production-Ready Terraform Module</h2> <div class="callout callout-tip"> <strong>💼 Production Grade:</strong> This is copy-paste ready for your infrastructure. Use this. Don't reinvent it. </div> <div class="code-block"> <div class="code-header">modules/aks_backup/main.tf</div> <pre><code class="language-hcl">terraform { required_providers { azurerm = { source = "hashicorp/azurerm" version = "~> 3.80" } } } variable "cluster_name" { type = string } variable "cluster_id" { type = string } variable "location" { type = string } variable "resource_group_name" { type = string } # ===== BACKUP VAULT ===== resource "azurerm_data_protection_backup_vault" "main" { name = "${var.cluster_name}-vault" resource_group_name = var.resource_group_name location = var.location datastore_type = "VaultStore" redundancy = "GeoRedundant" identity { type = "SystemAssigned" } tags = { Environment = "production" Purpose = "aks-cluster-backup" } } # ===== BACKUP POLICY ===== resource "azurerm_data_protection_backup_policy_kubernetes_cluster" "main" { name = "${var.cluster_name}-backup-policy" resource_group_name = var.resource_group_name vault_name = azurerm_data_protection_backup_vault.main.name backup_repeating_time_intervals = ["R/2026-01-01T00:00:00Z/PT4H"] default_retention_rule { lifetime = "P7D" } retention_rule { name = "weekly-retention" lifetime = "P4W" criteria { days_of_week = ["Sunday"] scheduled_backup_times = ["2026-01-01T02:00:00Z"] } } } # ===== TRUSTED ACCESS ===== resource "azurerm_role_assignment" "vault_aks_backup" { scope = var.cluster_id role_definition_name = "Kubernetes Service Cluster User Role" principal_id = azurerm_data_protection_backup_vault.main.identity[0].principal_id } # ===== OUTPUTS ===== output "backup_vault_id" { value = azurerm_data_protection_backup_vault.main.id }</code></pre> </div> <div class="code-block"> <div class="code-header">Usage: main.tf</div> <pre><code>module "aks_backup" { source = "./modules/aks_backup" cluster_name = azurerm_kubernetes_cluster.prod.name cluster_id = azurerm_kubernetes_cluster.prod.id location = azurerm_resource_group.main.location resource_group_name = azurerm_resource_group.main.name }</code></pre> </div> </section> <section class="decisions"> <h2>3 Critical Configuration Decisions</h2> <div class="decision-card"> <h3>Decision 1: When to Backup (Frequency Matters)</h3> <div class="comparison-table"> <div class="comparison-row"> <div class="comparison-item"> <strong>Every 4 Hours</strong> <p class="text-muted">Higher cost, lower RPO (4 hours max data loss)</p> </div> <div class="comparison-item"> <strong>Daily</strong> <p class="text-muted">Lower cost, higher RPO (24 hours max data loss)</p> </div> </div> </div> <div class="alert alert-info"> <strong>Right Answer:</strong> Match your business needs. Don't over-engineer costs. </div> </div> <div class="decision-card"> <h3>Decision 2: Redundancy Strategy (Safety vs Cost)</h3> <div class="comparison-table"> <div class="comparison-row"> <div class="comparison-item"> <strong>Production</strong> <p class="text-muted">GeoRedundant (survives regional Azure outages)</p> </div> <div class="comparison-item"> <strong>Development</strong> <p class="text-muted">LocallyRedundant (saves 30% on backups)</p> </div> </div> </div> </div> <div class="decision-card"> <h3>Decision 3: What NOT to Backup</h3> <div class="callout callout-warning"> <strong>Don't backup kube-system namespace.</strong> System components are ephemeral and reproducible. </div> <p><strong>Why:</strong> Saves 20-30% of snapshot costs. Backup only application data.</p> </div> </section> <section class="costs"> <h2>What Does It Actually Cost?</h2> <div class="pricing-table"> <table> <thead> <tr> <th>Component</th> <th>Monthly Cost</th> <th>Notes</th> </tr> </thead> <tbody> <tr> <td><strong>Backup Vault</strong></td> <td class="cost-primary">$50-100</td> <td>Flat rate per vault</td> </tr> <tr> <td><strong>Snapshots</strong></td> <td class="cost-primary">$0.05-0.15/GB</td> <td>Incremental + deduplication</td> </tr> <tr> <td><strong>Storage Logs</strong></td> <td class="cost-primary">$0.02-0.05/GB</td> <td>Minimal overhead</td> </tr> <tr class="total-row"> <td><strong>TOTAL</strong></td> <td class="cost-total"><strong>$200-500</strong></td> <td>Varies by cluster size</td> </tr> </tbody> </table> </div> <div class="cost-savings-box"> <h4>💰 How to Save Money</h4> <ul> <li><strong>Exclude system namespaces</strong> → 30% savings</li> <li><strong>Daily backups for dev</strong> → Lower costs</li> <li><strong>Shorter retention for dev</strong> → 7 days vs 30 days</li> <li><strong>LocallyRedundant for dev</strong> → 30% cost reduction</li> </ul> </div> </section> <section class="mistakes"> <h2>5 Mistakes That Will Hurt You</h2> <div class="mistake-box"> <div class="mistake-header"> <span class="mistake-icon">❌</span> <h3>Mistake 1: Never Testing Restores</h3> </div> <div class="mistake-content"> <p><strong>Impact:</strong> Corrupted backups discovered during actual disaster</p> <div class="alert alert-danger">Worst case: Complete data loss discovered only when needed</div> <p><strong>Fix:</strong> Do monthly restore tests to temporary cluster</p> </div> </div> <div class="mistake-box"> <div class="mistake-header"> <span class="mistake-icon">❌</span> <h3>Mistake 2: Backing Up Everything</h3> </div> <div class="mistake-content"> <p><strong>Impact:</strong> Snapshot costs spiral. Backups slow. Unnecessary data.</p> <p><strong>Fix:</strong> Exclude system namespaces explicitly</p> </div> </div> <div class="mistake-box"> <div class="mistake-header"> <span class="mistake-icon">❌</span> <h3>Mistake 3: Forgetting Workload Identity After Restore</h3> </div> <div class="mistake-content"> <p><strong>Impact:</strong> Apps can't authenticate. Silent failures everywhere.</p> <p><strong>Fix:</strong> Document OIDC issuer update in your runbook</p> </div> </div> <div class="mistake-box"> <div class="mistake-header"> <span class="mistake-icon">❌</span> <h3>Mistake 4: No Monitoring of Backup Jobs</h3> </div> <div class="mistake-content"> <p><strong>Impact:</strong> Backups fail silently for weeks</p> <p><strong>Fix:</strong> Set up Azure Monitor alerts for failed jobs</p> </div> </div> <div class="mistake-box"> <div class="mistake-header"> <span class="mistake-icon">❌</span> <h3>Mistake 5: One Vault for All Regions</h3> </div> <div class="mistake-content"> <p><strong>Impact:</strong> Regional Azure outage breaks both production and backups</p> <p><strong>Fix:</strong> One vault per cluster, use geo-redundancy</p> </div> </div> </section> <section class="runbook"> <h2>Disaster Recovery Runbook (Namespace Deletion)</h2> <div class="callout callout-important"> <strong>⏱️ Total Recovery Time: ~30 minutes</strong> </div> <div class="runbook-step"> <div class="step-header"> <span class="step-number">1</span> <h3>Assess (5 minutes)</h3> </div> <div class="code-block"> <pre><code>kubectl get ns | grep deleted-namespace az dataprotection backup-instance show \ --vault-name my-vault \ --resource-group my-rg \ --backup-instance-name my-backup</code></pre> </div> </div> <div class="runbook-step"> <div class="step-header"> <span class="step-number">2</span> <h3>Restore to Temporary Cluster (15 minutes)</h3> </div> <div class="code-block"> <pre><code>az aks create \ --resource-group my-rg \ --name temp-restore-cluster \ --node-count 2 az dataprotection job create \ --vault-name my-vault \ --resource-group my-rg \ --backup-instance-name my-backup-instance # Wait 5-10 minutes for restore to complete</code></pre> </div> </div> <div class="runbook-step"> <div class="step-header"> <span class="step-number">3</span> <h3>Extract and Reapply (10 minutes)</h3> </div> <div class="code-block"> <pre><code>kubectl config use-context temp-restore-cluster kubectl get all -n deleted-namespace -o yaml > namespace-backup.yaml kubectl config use-context prod-cluster kubectl apply -f namespace-backup.yaml</code></pre> </div> </div> <div class="runbook-step"> <div class="step-header"> <span class="step-number">4</span> <h3>Cleanup (2 minutes)</h3> </div> <div class="code-block"> <pre><code>az aks delete --name temp-restore-cluster --resource-group my-rg --yes</code></pre> </div> </div> </section> <section class="conclusion"> <h2>The Bottom Line</h2> <div class="conclusion-box"> <h3>🛡️ AKS backup isn't optional. It's insurance against catastrophe.</h3> <p>The teams that invest in proper backup strategy now are the ones who sleep soundly later. The teams that skip it? They're the ones rebuilding infrastructure at 3 AM.</p> </div> <h3>5 Things to Do Right Now</h3> <div class="action-list"> <div class="action-item"> <span class="action-number">1</span> <p>Define your RPO/RTO (Recovery Point Objective / Recovery Time Objective)</p> </div> <div class="action-item"> <span class="action-number">2</span> <p>Deploy this Terraform module to production TODAY</p> </div> <div class="action-item"> <span class="action-number">3</span> <p>Schedule your first restore test</p> </div> <div class="action-item"> <span class="action-number">4</span> <p>Set up Azure Monitor alerts for failed backups</p> </div> <div class="action-item"> <span class="action-number">5</span> <p>Document your Workload Identity OIDC update procedure</p> </div> </div> <div class="alert alert-success"> <strong>Remember:</strong> The investment pays dividends the moment disaster strikes. Don't be the team rebuilding from scratch. </div> </section> </div>
← Back to Blog