devops

Build a Helm Release Health Agent: Fix Stuck pending-upgrade Releases, Failed Hooks, and 'another operation is in progress' Safely

Build a Helm release health agent that finds stuck pending-upgrade releases, failed hooks and 'another operation is in progress' errors, then rolls back safely.

September 29, 2026·11 min read·
#ai#devops#kubernetes#automation#reliability#gitops#cicd

What This Agent Does

A Helm release health agent scans every Helm release in the cluster, finds the ones whose latest revision is stuck in pending-upgrade, pending-install, or failed, works out which of five failure classes it belongs to, and either rolls the release back to the last revision that actually deployed or hands a human a precise diagnosis. The scan and the fix are deterministic code. The LLM only classifies a stuck release from the evidence and picks from a short list of actions the executor already knows how to do safely. It cannot run helm uninstall, cannot pass --force, and cannot delete any Secret that is not a Helm release record in a pending state.

The reason to build it is the error every CI pipeline eventually prints:

Error: UPGRADE FAILED: another operation (install/upgrade/rollback) is in progress

Helm has no lock file. That message means the newest release record has a pending-* status, and Helm refuses to start a second operation on top of it. Nine times out of ten, no operation is in progress at all. A previous helm upgrade --wait was killed by a CI timeout and left its record behind. The fix takes one command, but only if you pick the right revision and the right command, and that is exactly where people break production at 2 a.m.

Where Helm Keeps Release State

Helm 3 stores each revision of a release as a Secret of type helm.sh/release.v1 in the release namespace, named sh.helm.release.v1.<name>.v<revision>. The labels are the fast path for an inventory, because they carry the status without decoding anything:

kubectl get secrets -A -l owner=helm \
  -o custom-columns='NS:.metadata.namespace,RELEASE:.metadata.labels.name,REV:.metadata.labels.version,STATUS:.metadata.labels.status,CREATED:.metadata.creationTimestamp' \
  | grep -E 'pending-|failed'
prod    api        13   pending-upgrade   2026-09-29T02:11:40Z
prod    worker     41   failed            2026-09-29T02:14:03Z
staging billing    7    pending-install   2026-09-27T16:02:19Z

The release data field is base64 wrapped around a gzipped JSON document, and the API server base64-encodes it once more, so it decodes twice. The info.description inside is the single most useful string in the whole system: Helm writes the actual error into it.

kubectl get secret -n prod sh.helm.release.v1.worker.v41 -o jsonpath='{.data.release}' \
  | base64 -d | base64 -d | gunzip \
  | jq '{version, status: .info.status, description: .info.description, chart: .chart.metadata.version}'
{
  "version": 41,
  "status": "failed",
  "description": "Upgrade \"worker\" failed: pre-upgrade hooks failed: job failed: BackoffLimitExceeded",
  "chart": "1.18.2"
}

Two things about how Helm behaves matter for every decision below. First, when the latest revision is failed, the next helm upgrade diffs against the last revision whose status is deployed, not against the failed one, so a plain re-run is usually safe. Second, when the latest revision is pending-*, helm upgrade refuses to run at all, but helm rollback does not check for pending state, which is why rollback is the documented way out of a stale lock.

The Five Failure Classes

ClassWhat info.description or the CI log saysSafe action
Stale lockStatus is pending-upgrade or pending-install, record is older than any --timeout you use, no helm process is aliveRoll back to the last deployed revision
Wait timeouttimed out waiting for the conditionDiagnose the pods, then roll back or re-run
Failed hookpre-upgrade hooks failed: job failed: BackoffLimitExceededRead the hook Job logs, fix the migration, re-run
Ownership conflictrendered manifests contain a resource that already exists ... invalid ownership metadataAdopt the resource with Helm's labels via a PR
Immutable fieldcannot patch ... field is immutableEscalate; the fix needs a delete-and-recreate

The first row is the one worth automating end to end. The stale-lock case is mechanical, common, and the fix is exactly one command. The other four need a human to see a diagnosis, and the agent's job is to make that diagnosis complete on the first page.

Step 1: The Inventory Is Code

The scanner lists release Secrets, keeps only the latest revision per release, and decodes the payload only for releases that look stuck. It also detects releases owned by Flux, because helm-controller labels its release Secrets with helm.toolkit.fluxcd.io/name, and touching those by hand just triggers a reconcile that undoes your fix. Argo CD renders Helm charts as plain manifests and never writes a release Secret, so Argo-managed apps never appear here at all. If you run Argo, the Argo CD MCP server is the right tool for a stuck sync.

# inventory.py — read-only, runs in-cluster
import base64, gzip, json
from datetime import datetime, timezone
from kubernetes import client, config

config.load_incluster_config()
core = client.CoreV1Api()
PENDING = {"pending-install", "pending-upgrade", "pending-rollback", "uninstalling"}
STALE_AFTER_S = 1800   # 2x the longest --timeout used in CI

def decode(secret) -> dict:
    raw = base64.b64decode(base64.b64decode(secret.data["release"]))
    return json.loads(gzip.decompress(raw))

def release_inventory() -> list[dict]:
    latest = {}
    for s in core.list_secret_for_all_namespaces(label_selector="owner=helm").items:
        lb = s.metadata.labels
        key = (s.metadata.namespace, lb["name"])
        rev = int(lb["version"])
        if key in latest and rev <= latest[key]["revision"]:
            continue
        age = (datetime.now(timezone.utc) - s.metadata.creation_timestamp).total_seconds()
        latest[key] = {
            "release": f"{key[0]}/{key[1]}",
            "revision": rev,
            "status": lb["status"],
            "age_s": int(age),
            "flux_managed": "helm.toolkit.fluxcd.io/name" in lb,
            "secret": s,
        }
    stuck = []
    for r in latest.values():
        if r["status"] == "deployed":
            continue
        if r["status"] in PENDING and r["age_s"] < STALE_AFTER_S:
            continue   # a real helm process may still be running; check again later
        info = decode(r.pop("secret"))["info"]
        r["description"] = info["description"]
        r["last_deployed_revision"] = last_deployed(r["release"])
        stuck.append(r)
    return stuck

def last_deployed(release: str) -> int | None:
    ns, name = release.split("/")
    revs = core.list_namespaced_secret(ns, label_selector=f"owner=helm,name={name},status=deployed").items
    return max((int(s.metadata.labels["version"]) for s in revs), default=None)

Two details are deliberate. A pending record younger than the stale threshold is skipped entirely, because an upgrade with a long --timeout can legitimately sit in pending-upgrade for 15 minutes. And the last deployed revision comes from a label query, never from arithmetic like "current minus one". A release that failed three times in a row has three failed records above its last deployed one, and rolling back to revision - 1 would roll back to another failure.

Step 2: Gather Evidence Per Stuck Release

For each stuck release the agent pulls the same four things an engineer would, and puts them in a fixed-shape bundle for the model. The commands run through the read-only Kubernetes MCP server so the model never gets a shell:

helm history worker -n prod -o json | jq '.[-4:]'
helm get hooks worker -n prod | grep -E '^kind:|helm.sh/hook'
kubectl get jobs -n prod -l app.kubernetes.io/instance=worker \
  -o custom-columns='NAME:.metadata.name,SUCCEEDED:.status.succeeded,FAILED:.status.failed'
kubectl logs -n prod job/worker-db-migrate --tail=40

The history output is where the wait-timeout and stale-lock cases become obvious:

[
  {"revision": 39, "status": "superseded", "description": "Upgrade complete"},
  {"revision": 40, "status": "deployed",   "description": "Upgrade complete"},
  {"revision": 41, "status": "failed",     "description": "Upgrade \"worker\" failed: pre-upgrade hooks failed: job failed: BackoffLimitExceeded"}
]

For a wait timeout, the bundle also includes the pods behind the release and their waiting reasons. A --wait that timed out is almost always a CrashLoopBackOff or a readiness probe that never passed, and the agent should say which, not just "timed out".

Hook Jobs are only inspectable if the chart keeps them. A hook annotated with helm.sh/hook-delete-policy: hook-succeeded leaves failed Jobs in place, which is what you want. A policy of before-hook-creation,hook-succeeded,hook-failed deletes the evidence, and the agent reports that explicitly so someone fixes the chart rather than wondering where the logs went. The Helm chart best practices post covers the hook conventions worth standardising on.

Step 3: The Model Classifies, the Executor Acts

The model gets one stuck release at a time with the evidence bundle and must return this object. The executor validates it against a JSON schema before it does anything:

{
  "release": "prod/worker",
  "diagnosis": "failed_hook",
  "evidence": [
    "info.description: pre-upgrade hooks failed: job failed: BackoffLimitExceeded",
    "job worker-db-migrate FAILED=3, last log line: relation \"orders_v2\" already exists"
  ],
  "action": "none_escalate",
  "target_revision": null,
  "summary_for_owner": "Migration 0142 is not idempotent; it fails on re-run because orders_v2 exists. Release is 'failed' but revision 40 is still serving. Fix the migration and re-run the pipeline; no rollback needed.",
  "confidence": "high"
}

The diagnosis enum is the five classes above plus flux_managed and unknown. The action enum is rollback_to_last_deployed, clear_stale_lock, rerun_upgrade, or none_escalate. The system prompt is short and mostly rules:

You classify a stuck Helm release. Use only the evidence bundle.
- A pending-* status older than the stale threshold with no live helm process
  is a stale lock. The correct action is rollback_to_last_deployed.
- A failed hook or an immutable-field error is never fixed by rollback. Escalate
  with the exact log line that explains the failure.
- If last_deployed_revision is null, you cannot roll back. Escalate.
- If flux_managed is true, the only action is none_escalate; say to run
  `flux reconcile hr <name> -n <ns>` instead.
- Every evidence entry must quote a string from the bundle verbatim.

The executor is where the safety lives, and it is plain Python, not a prompt:

def execute(decision: dict, rel: dict):
    if rel["flux_managed"] or decision["action"] == "none_escalate":
        return notify_owner(decision)
    if decision["action"] == "rollback_to_last_deployed":
        target = rel["last_deployed_revision"]
        assert target is not None and target < rel["revision"]
        assert status_of(rel["release"], target) == "deployed"
        return run(["helm", "rollback", name, str(target), "-n", ns,
                    "--wait", "--timeout", "5m", "--cleanup-on-fail"])
    if decision["action"] == "clear_stale_lock":
        # only when rollback is impossible: pending-install with nothing deployed
        assert rel["status"] == "pending-install" and rel["last_deployed_revision"] is None
        return core.delete_namespaced_secret(
            f"sh.helm.release.v1.{name}.v{rel['revision']}", ns)

Note the asymmetry. rollback_to_last_deployed is the default fix for a stale lock because it is a normal Helm operation: it writes a new revision with the manifest of the target and marks the pending record superseded. Deleting the pending Secret by hand also works, and plenty of runbooks recommend it, but the executor allows it only for a pending-install with no deployed revision to fall back to. Deleting the wrong Secret, in particular the deployed one, makes Helm forget the release exists while the workloads keep running, and the next upgrade fails with has no deployed releases. The assertion on the status label is the guard against that.

Step 4: Approval and Blast Radius

Rollback re-applies the last deployed manifest with --wait, which is a real deployment. In a prod namespace that goes through the same human-in-the-loop approval gate as any other write, with the message showing the diff between the target revision and what is currently running:

helm get manifest worker -n prod --revision 40 > /tmp/target.yaml
helm get manifest worker -n prod --revision 41 > /tmp/current.yaml
diff -u /tmp/current.yaml /tmp/target.yaml | grep -E '^[+-]\s+image:'

If the only lines that differ are image tags, the reviewer knows the rollback moves one version. If a ConfigMap or a Secret name differs, the reviewer knows to look harder. In non-prod namespaces the agent rolls back on its own, and every action lands in the audit trail with the release, the revisions involved, the decision object, and the pipeline run that left the lock behind.

The ServiceAccount needs get and list on Secrets cluster-wide to read release records, plus delete on Secrets restricted by a resource-name prefix, which RBAC cannot express. The practical answer is to grant delete only in namespaces where clear_stale_lock is permitted at all, and to keep the executor's assertion as the second line of defence. The RBAC for AI agents post has the Role template.

Fixing the Pipeline That Causes It

The agent cleans up after the failure. The failure itself comes from a CI step that runs helm upgrade --wait with a --timeout longer than the job's own timeout, so the runner kills helm mid-flight. Three changes make stale locks rare:

helm upgrade --install worker ./chart -n prod \
  --atomic --timeout 8m --history-max 15 \
  --set image.tag="$GIT_SHA"

--atomic rolls back automatically if the upgrade fails, so a wait timeout never leaves a failed record. The job-level timeout must be longer than the Helm --timeout plus a rollback, so the runner never kills the process. And --history-max keeps enough deployed revisions that the last good one is still there to roll back to when something does slip through. Add the agent as the safety net for the cases where a runner dies anyway, and the "another operation is in progress" error stops being a page.

Honest Limits

The agent cannot tell whether a helm process is genuinely still running from outside the cluster, which is why the stale threshold exists and why it should be generous. It cannot fix an immutable-field error, because the fix is to delete the Deployment or change its name, and both are release-owner decisions. And it does not help at all with Argo CD, which has its own lock semantics and its own tooling. What it does is turn the most common Helm failure in CI into a diagnosis that quotes the exact log line, and a rollback to a revision it has verified was actually deployed.

#ai#devops#kubernetes#automation#reliability#gitops#cicd
D
DevToCashAuthor

Senior DevOps/SRE Engineer · 10+ years · Professional Trader (IDX, Crypto, US Equities)

I write about real infrastructure patterns and trading strategies I use in production and in live markets. No courses, no affiliate hype — just documentation of what actually works.

More about me →