sre

Build a Backup Verification Agent: Velero Restore Drills That Prove Your Kubernetes Backups Actually Restore

Build a Velero backup verification agent for Kubernetes: catch PartiallyFailed backups, stale schedules and run restore drills that prove data comes back.

October 1, 2026·10 min read·
#ai#sre#kubernetes#reliability#automation

What This Agent Does

A backup verification agent answers one question every morning: if the prod-shop namespace disappeared right now, would it come back? It reads Velero's Schedules, Backups, and BackupStorageLocations, flags schedules whose last successful backup is older than their cron interval, triages every PartiallyFailed backup into a named cause, and once a week restores a real backup into a throwaway namespace and runs a check against the restored data. Deterministic code does the inventory and the restore. The LLM gets the error strings and a short list of allowed diagnoses. The one write action, starting a restore drill, is gated behind an approval and hard-wired to a namespace prefix that production never uses.

The reason to build it is that velero backup get showing a column of Completed is not evidence. A backup that completed with the wrong include list, with volumes skipped, or against a database that was mid-write is still Completed. Only a restore tells you what you have.

Why "Completed" Lies

Velero reports each backup with a phase, and the phases that look fine hide most of the problems.

What you seeWhat it can mean
Completed, 0 errorsResources backed up, but PVCs were silently excluded because no snapshot class matched, or defaultVolumesToFsBackup was off and the pods had no opt-in annotation
PartiallyFailed, 3 errorsOften one node-agent timeout on a large volume, which means the application data is missing while all 400 ConfigMaps are fine
Schedule Enabled, last backup 9 days agoThe controller restarted, the cron is 0 2 * * *, and nobody looked at status.lastBackup
BSL AvailableValidated at the last check, which may be 30 minutes old, and validation only tests listing the bucket, not writing to it
Completed restoreResources created; pods may be CrashLoopBackOff because a Secret was excluded or a PV never attached

The agent treats all five as findings, and the fifth is why it runs a verifier Job after every drill rather than trusting the restore phase. If the restored pods go to FailedMount or FailedAttachVolume, the causes are the same ones in the FailedMount troubleshooting guide, usually a StorageClass or CSI driver mismatch between the snapshot and the target.

Step 1: Inventory Is Code

The schedule check needs no model. It parses each Schedule's cron, computes the expected gap, and compares it to status.lastBackup:

# inventory.py — read-only against the velero namespace
from datetime import datetime, timezone, timedelta
from croniter import croniter
from kubernetes import client, config

config.load_incluster_config()
crd = client.CustomObjectsApi()
NS = "velero"

def get(kind):
    return crd.list_namespaced_custom_object("velero.io", "v1", NS, kind)["items"]

def stale_schedules(grace: float = 1.5) -> list[dict]:
    now = datetime.now(timezone.utc)
    out = []
    for s in get("schedules"):
        if s["spec"].get("paused"):
            continue
        cron = s["spec"]["schedule"]
        interval = croniter(cron, now).get_next(datetime) - croniter(cron, now).get_prev(datetime)
        last = s.get("status", {}).get("lastBackup")
        last_dt = datetime.fromisoformat(last.replace("Z", "+00:00")) if last else None
        if last_dt is None or now - last_dt > interval * grace:
            out.append({"schedule": s["metadata"]["name"], "cron": cron,
                        "last_backup": last, "expected_every": str(interval)})
    return out

def problem_backups(hours: int = 48) -> list[dict]:
    cutoff = datetime.now(timezone.utc) - timedelta(hours=hours)
    out = []
    for b in get("backups"):
        st = b.get("status", {})
        started = st.get("startTimestamp")
        if not started or datetime.fromisoformat(started.replace("Z", "+00:00")) < cutoff:
            continue
        if st.get("phase") in ("Completed",) and not st.get("errors"):
            continue
        out.append({"backup": b["metadata"]["name"],
                    "schedule": b["metadata"].get("labels", {}).get("velero.io/schedule-name"),
                    "phase": st.get("phase"), "errors": st.get("errors", 0),
                    "warnings": st.get("warnings", 0),
                    "items_backed_up": st.get("progress", {}).get("itemsBackedUp"),
                    "total_items": st.get("progress", {}).get("totalItems"),
                    "volumes_fs_backup": st.get("backupItemOperationsAttempted", 0)})
    return out

The same loop checks backupstoragelocations for status.phase != Available and alerts outright, because nothing downstream matters if the bucket is unreachable. For clusters that already scrape Velero's metrics, the read-only Prometheus MCP server can answer the staleness question in one query:

# seconds since each schedule last succeeded; alert above 1.5x the interval
time() - velero_backup_last_successful_timestamp{schedule!=""}

# schedules whose most recent backup failed or partially failed
velero_backup_last_status == 0

The CRD walk is still worth having because the metric only exists for schedules that have ever succeeded. A schedule that has never produced a backup is invisible in Prometheus and is the most dangerous kind.

Step 2: Give the Model Errors, Not Objects

For each PartiallyFailed backup the wrapper pulls the error section of the describe output and a bounded grep of the backup log. The full log for a large backup can be 50 MB; the model gets 40 lines.

velero backup describe prod-shop-20261001020012 --details | sed -n '/^Errors:/,/^Warnings:/p'
velero backup logs prod-shop-20261001020012 | grep -E 'level=error' | tail -n 40

Typical output:

Errors:
  Velero:    <none>
  Cluster:   <none>
  Namespaces:
    prod-shop:  error backing up item: timed out waiting for all PodVolumeBackups to complete
time="2026-10-01T02:31:44Z" level=error msg="Error backing up item" backup=velero/prod-shop-20261001020012 error="timed out waiting for all PodVolumeBackups to complete" name=postgres-0 namespace=prod-shop resource=pods

The tool the model calls to see this returns a fixed shape, and the schema is the whole contract:

{
  "name": "describe_backup_errors",
  "description": "Return the error section and up to 40 error-level log lines for one Velero backup. Read-only.",
  "input_schema": {
    "type": "object",
    "properties": {
      "backup": {"type": "string", "pattern": "^[a-z0-9-]{1,63}$"}
    },
    "required": ["backup"]
  }
}

The system prompt limits the model to a closed set of diagnoses, each tied to an action the wrapper knows how to execute or to a human it knows how to page:

You triage Velero backup failures. For each backup, pick exactly one cause:
  fs_backup_timeout      - PodVolumeBackup/DataUpload timed out; large volume or node-agent pressure
  snapshot_plugin_error  - CSI/cloud snapshot failed (VolumeSnapshotClass, quota, permissions)
  bsl_unavailable        - object store unreachable or credentials rejected
  excluded_volume        - volume skipped: no fs-backup opt-in and no matching snapshot class
  transient_item_error   - errors only on Events, EndpointSlices, or Leases; data is intact
  unknown                - none of the above; include the exact error string
Report the evidence line that decided it. Never propose kubectl or velero commands.

The transient_item_error class matters more than it looks. Velero backing up a namespace with churny EndpointSlices produces a PartiallyFailed phase with "the object has been modified" errors on resources nobody would ever restore. Without that class, the agent would wake someone up for noise, which is the failure mode the alert tuning post is about. With it, the wrapper downgrades the finding to a warning and recommends adding those resources to excludedResources on the Schedule.

Step 3: The Restore Drill Is the Only Write, and It Is Boxed In

The drill restores one namespace from its latest Completed backup into a new namespace with a reserved prefix, under three constraints the model cannot change:

  1. The target namespace must match ^drill-[a-z0-9-]+-[0-9]{8}$ and must not exist yet.
  2. The restore excludes resources that would collide with or cost like production: ingresses, services of type LoadBalancer, horizontalpodautoscalers, and anything cluster-scoped.
  3. Every Deployment and StatefulSet is restored with zero replicas, so restored pods never start, never connect to production databases with the restored Secrets, and never consume capacity beyond the PVCs.

The third constraint uses a Velero resource modifier, which is a ConfigMap in the velero namespace:

apiVersion: v1
kind: ConfigMap
metadata:
  name: drill-modifiers
  namespace: velero
data:
  modifiers.yaml: |
    version: v1
    resourceModifierRules:
    - conditions:
        groupResource: deployments.apps
      patches:
      - operation: replace
        path: "/spec/replicas"
        value: "0"
    - conditions:
        groupResource: statefulsets.apps
      patches:
      - operation: replace
        path: "/spec/replicas"
        value: "0"

The wrapper then builds the restore. The model's input is only the source namespace, and even that is validated against an allowlist:

import subprocess, re
from datetime import date

DRILL_SOURCES = {"prod-shop", "prod-billing"}   # namespaces with a drill owner

def start_restore_drill(source_ns: str, backup: str) -> str:
    if source_ns not in DRILL_SOURCES or not re.fullmatch(r"[a-z0-9-]{1,63}", backup):
        raise ValueError("refused")
    target = f"drill-{source_ns}-{date.today():%Y%m%d}"
    name = f"{target}-restore"
    subprocess.run(["velero", "restore", "create", name,
        "--from-backup", backup,
        "--include-namespaces", source_ns,
        "--namespace-mappings", f"{source_ns}:{target}",
        "--exclude-resources", "ingresses,horizontalpodautoscalers,poddisruptionbudgets",
        "--include-cluster-resources=false",
        "--resource-modifier-configmap", "drill-modifiers",
        "--labels", f"drill=true,source={source_ns}",
        "--wait"], check=True, timeout=1800)
    return name

LoadBalancer Services need one more step because Velero has no type-based exclude. The wrapper adds a second modifier rule that replaces /spec/type with ClusterIP for services, which stops the restore from provisioning a cloud load balancer that would sit idle at a few dollars a day and, worse, carry production DNS annotations.

Starting the drill goes through the same approval gate as every write in this series. The request a human sees is short: source namespace, backup name and age, target namespace, and the estimated PVC storage the drill will allocate. The approval gates post covers the mechanics. In practice, after the first month the approval becomes a weekly rubber stamp, and that is fine. The gate exists so that a prompt-injected error message in a backup log cannot trigger a restore into a namespace of the attacker's choosing.

Step 4: Verify the Data, Not the Phase

After the restore reports Completed, the wrapper runs a verifier Job in the drill namespace that mounts the restored PVC and checks it. For Postgres this is the useful minimum:

apiVersion: batch/v1
kind: Job
metadata:
  name: verify-postgres
  namespace: drill-prod-shop-20261001
spec:
  backoffLimit: 0
  ttlSecondsAfterFinished: 3600
  template:
    spec:
      restartPolicy: Never
      containers:
      - name: verify
        image: postgres:16
        command: ["sh", "-c"]
        args:
        - |
          set -e
          pg_controldata /var/lib/postgresql/data | grep -E 'Database cluster state|Latest checkpoint time'
          test "$(find /var/lib/postgresql/data/base -type f | wc -l)" -gt 100
        volumeMounts:
        - name: data
          mountPath: /var/lib/postgresql/data
      volumes:
      - name: data
        persistentVolumeClaim:
          claimName: data-postgres-0

Two lines of output decide the finding. Database cluster state: in production means the snapshot caught Postgres mid-run, and the restore will work but will replay WAL on first start; shut down means a clean snapshot. A checkpoint time more than a few minutes before the backup's startTimestamp means the pre-backup hook did not flush. The file-count check is deliberately crude. It catches the case that bites teams most: a PVC that restored empty because the snapshot class was wrong and Velero created a fresh volume.

The model sees the Job's log and the backup record, and produces the one-paragraph finding that goes to the owning team. The wrapper deletes the drill namespace after the Job finishes, which also releases the restored PVCs. On a cluster where storage is tight, the capacity planning agent should know about the drill schedule, because a 400 GB restore that lands on the wrong day can push a real workload into Pending.

RBAC and Blast Radius

The agent's ServiceAccount gets get and list on Velero CRDs in the velero namespace, create on restores.velero.io only, and full access inside namespaces matching the drill- prefix. It cannot create Backups, so it can never overwrite a retention policy or flood the bucket, and it cannot delete Backups, which is the one Velero action that destroys recovery options. The velero CLI inside the container is invoked with a kubeconfig bound to that ServiceAccount, not the cluster-admin one the Velero server uses. If you have followed the RBAC for AI agents post, this is one more Role in the same pattern.

Honest Limits

The drill verifies that data restores on this cluster, with this CSI driver, into this StorageClass. It says nothing about restoring to a different region after the primary is gone, which needs a periodic drill on a second cluster that this agent can schedule but not judge from inside the first. The verifier Jobs are per-application and someone has to write them; a generic file count is better than nothing and worse than a real consistency check. And the LLM triage is only as good as the error strings Velero emits, which for CSI data movement failures sometimes amount to "context deadline exceeded" and nothing more. When the model returns unknown, the wrapper pages a human with the raw log instead of guessing, which is the right failure mode for a system whose entire job is to be trusted on the worst day of the year.

#ai#sre#kubernetes#reliability#automation
D
DevToCashAuthor

Senior DevOps/SRE Engineer · 10+ years · Professional Trader (IDX, Crypto, US Equities)

I write about real infrastructure patterns and trading strategies I use in production and in live markets. No courses, no affiliate hype — just documentation of what actually works.

More about me →