sre

Build a TLS Certificate Expiry Agent: Catch cert-manager Renewal Failures, Unmanaged Secrets, and Stale Certs Before the Outage

Build a TLS certificate expiry agent for Kubernetes: diagnose cert-manager renewal failures, unmanaged TLS secrets and stale served certs before they expire.

September 28, 2026·10 min read·
#ai#sre#kubernetes#reliability#automation#security

What This Agent Does

A TLS certificate expiry agent finds every certificate in and around your cluster that will expire soon, works out why it has not renewed, and either fixes it or hands a specific owner a specific task. The inventory is deterministic code: it reads cert-manager's metrics, parses every kubernetes.io/tls Secret, and connects to each Ingress host to see which certificate is actually being served. The LLM only sees the shortlist of certificates that are inside the renewal window and still not renewed, and for each one it picks from a small set of diagnoses and remediations. Private keys never leave the process, and the agent cannot create, delete, or edit a certificate by hand. It can re-trigger a cert-manager renewal, restart a workload that is serving a stale cert, or open a pull request.

The reason to build it is that "we run cert-manager" is not the same as "our certificates renew". Every expiry outage I have seen in the last three years happened on a cluster that had cert-manager installed.

Why Certificates Still Expire With cert-manager Installed

cert-manager renews a Certificate at two thirds of its lifetime by default, so a 90-day Let's Encrypt cert enters its renewal window 30 days before expiry. That leaves a long window in which a renewal can be failing quietly. The failures fall into five classes, and each has a different fix.

ClassHow it shows upWho fixes it
ACME challenge failingCertificate has Ready=False, a Challenge sits in pending with a reason stringUsually an Ingress or DNS config fix in Git
Issuer broken or rate-limitedOrder is errored, message mentions urn:ietf:params:acme:error:rateLimited or the issuer is Ready=FalseWait, or change the issuer
Unmanaged SecretA kubernetes.io/tls Secret with no cert-manager.io/certificate-name annotation, created by hand or by a Helm chart, expiring on a date nobody wrote downThe owning team, ideally by adopting cert-manager
Renewed but not reloadedThe Secret holds a fresh cert, the pod is still serving the old one because it read the file at startupA rollout restart
Outside cert-manager's viewWebhook caBundles, kubelet serving certs, a load balancer terminating TLS with its own certDifferent tooling per case; the agent reports, never touches

The first two classes are visible in cert-manager's own objects. The third and fourth are not, and they are the ones that page you, because nothing was watching them. A stale cert on an ingress backend also produces the same symptoms as a bad upstream, so if you are chasing intermittent errors at the edge, the ingress-nginx 502 guide covers the non-TLS causes and this agent covers the TLS ones.

Step 1: The Inventory Is Code, Not a Prompt

Three sources, merged by Secret name. The first is cert-manager's metrics, which the read-only Prometheus MCP server can answer:

# Certificates expiring within 14 days
(certmanager_certificate_expiration_timestamp_seconds - time()) / 86400 < 14

# Certificates cert-manager itself says are not ready
certmanager_certificate_ready_status{condition="False"} == 1

Metrics only cover certificates cert-manager manages, so the second source walks every TLS Secret and parses the leaf certificate locally:

# inventory.py — deterministic, read-only, keys never leave this process
import base64, ssl, socket
from datetime import datetime, timezone
from cryptography import x509
from kubernetes import client, config

config.load_incluster_config()
core, net = client.CoreV1Api(), client.NetworkingV1Api()

def parse_leaf(pem: bytes) -> x509.Certificate:
    return x509.load_pem_x509_certificate(pem.split(b"-----END CERTIFICATE-----")[0]
                                          + b"-----END CERTIFICATE-----\n")

def secret_inventory() -> dict:
    out = {}
    for s in core.list_secret_for_all_namespaces(field_selector="type=kubernetes.io/tls").items:
        crt = s.data.get("tls.crt")
        if not crt:
            continue
        leaf = parse_leaf(base64.b64decode(crt))
        ann = s.metadata.annotations or {}
        key = f"{s.metadata.namespace}/{s.metadata.name}"
        out[key] = {
            "secret": key,
            "managed_by": ann.get("cert-manager.io/certificate-name"),   # None = unmanaged
            "issuer": ann.get("cert-manager.io/issuer-name"),
            "not_after": leaf.not_valid_after_utc.isoformat(),
            "days_left": (leaf.not_valid_after_utc - datetime.now(timezone.utc)).days,
            "serial": format(leaf.serial_number, "x"),
            "sans": [n.value for n in leaf.extensions.get_extension_for_class(
                x509.SubjectAlternativeName).value],
        }
    return out

The tls.key field is deliberately never read. The agent's ServiceAccount needs get and list on Secrets to do this at all, which is the most sensitive permission in this series, so the secrets-management post applies in full: the model receives the dictionary above, never the Secret object.

The third source connects to each Ingress host and compares the served serial to the stored one. This is the only way to catch class four:

def served_serial(host: str, port: int = 443) -> str | None:
    ctx = ssl.create_default_context()
    ctx.check_hostname, ctx.verify_mode = False, ssl.CERT_NONE   # we want the cert, not validation
    try:
        with socket.create_connection((host, port), timeout=5) as sock:
            with ctx.wrap_socket(sock, server_hostname=host) as tls:
                der = tls.getpeercert(binary_form=True)
        return format(x509.load_der_x509_certificate(der).serial_number, "x")
    except (OSError, ssl.SSLError):
        return None

def ingress_hosts() -> dict:
    hosts = {}
    for ing in net.list_ingress_for_all_namespaces().items:
        for tls in ing.spec.tls or []:
            for h in tls.hosts or []:
                hosts[h] = f"{ing.metadata.namespace}/{tls.secret_name}"
    return hosts

Merging the three gives a per-certificate record: days left, whether cert-manager owns it, whether cert-manager says it is ready, and whether what is served matches what is stored. Everything with more than 21 days left and a matching served serial is dropped. What remains is the shortlist, and on most clusters it is under ten items.

For cert-manager-managed items on the shortlist, the wrapper also pulls the chain of objects that explains the failure, because the reason string lives at the bottom of it:

kubectl get certificate,certificaterequest,order,challenge -n shop -o wide
NAME                                  READY   SECRET          AGE
certificate.cert-manager.io/shop-tls  False   shop-tls        88d

NAME                                          STATE     DOMAIN         REASON
challenge.acme.cert-manager.io/shop-tls-...   pending   shop.example   Waiting for HTTP-01 challenge propagation: wrong status code '404', expected '200'

That reason string is the single most useful input the model gets, and the wrapper passes it verbatim.

Step 2: The Diagnosis Call

One model call per shortlisted certificate, with a forced tool schema so the answer is structured:

CERT_TOOL = {
    "name": "diagnose_certificate",
    "description": "Explain why one certificate has not renewed and choose a remediation.",
    "input_schema": {
        "type": "object",
        "properties": {
            "diagnosis": {"enum": [
                "acme_http01_unreachable", "acme_dns01_not_propagated",
                "acme_rate_limited", "issuer_not_ready",
                "unmanaged_secret_expiring", "renewed_not_reloaded",
                "served_cert_is_external", "unknown"]},
            "action": {"enum": [
                "retrigger_renewal", "rollout_restart",
                "propose_ingress_fix_pr", "propose_adopt_cert_manager_pr",
                "wait_for_rate_limit", "ask_owner", "report_only"]},
            "workload": {"type": "string",
                "description": "namespace/kind/name to restart. Only for rollout_restart."},
            "evidence": {"type": "string",
                "description": "2-3 sentences quoting the challenge or order reason, "
                               "the days-left figure, and the served-vs-stored serial."},
            "risk": {"type": "string",
                "description": "What breaks if this action is wrong, and why the "
                               "safer alternative was not chosen."},
        },
        "required": ["diagnosis", "action", "evidence", "risk"],
    },
}

SYSTEM = (
    "You diagnose TLS certificate renewals for an SRE team. A 404 or connection "
    "refused on an HTTP-01 challenge is an Ingress routing problem: the "
    "/.well-known/acme-challenge/ path is not reaching the cert-manager solver "
    "pod. Choose propose_ingress_fix_pr. A DNS-01 'not yet propagated' reason "
    "under two hours old is normal; choose report_only. A rateLimited order is "
    "never fixed by retrying: choose wait_for_rate_limit and say when the window "
    "clears. If the Secret is fresh but the served serial is old, the pod read the "
    "cert at startup: choose rollout_restart and name the workload. An unmanaged "
    "Secret is always ask_owner or propose_adopt_cert_manager_pr, never "
    "retrigger_renewal. If the served certificate is not in any Secret at all, "
    "TLS terminates outside the cluster: served_cert_is_external, report_only. "
    "Use retrigger_renewal only when the Certificate is Ready=True, inside its "
    "renewal window, and no Order exists for it."
)

The last rule in the prompt matters. cmctl renew is the tempting hammer, and on a certificate with a failing challenge it does nothing except create another failed Order, which counts against the Let's Encrypt limit of five failed validations per hostname per hour. The agent is allowed to swing it only in the one situation where it helps: cert-manager simply has not got round to renewing yet, usually because the controller was down or restarted mid-cycle.

Step 3: Executing With Hard Limits

The RBAC, built the way the least-privilege ServiceAccount post lays out, encodes what the agent cannot do rather than trusting the prompt:

rules:
  - apiGroups: [""]
    resources: ["secrets"]
    verbs: ["get", "list"]                        # parse tls.crt; tls.key is never read
  - apiGroups: ["networking.k8s.io"]
    resources: ["ingresses"]
    verbs: ["get", "list"]
  - apiGroups: ["cert-manager.io"]
    resources: ["certificates", "certificaterequests", "issuers", "clusterissuers"]
    verbs: ["get", "list"]
  - apiGroups: ["cert-manager.io"]
    resources: ["certificates/status"]
    verbs: ["update"]                             # what cmctl renew needs, nothing more
  - apiGroups: ["acme.cert-manager.io"]
    resources: ["orders", "challenges"]
    verbs: ["get", "list"]
  - apiGroups: ["apps"]
    resources: ["deployments", "statefulsets", "daemonsets"]
    verbs: ["get", "patch"]                       # rollout restart annotation only

No create or delete on Secrets, so the agent cannot paste a certificate in by hand or wipe one that cert-manager will regenerate. No write on certificates, issuers, or ingresses, so every configuration change becomes a pull request through the GitOps-for-agents path.

The executor adds three rules RBAC cannot express:

  • rollout_restart is gated by days left. Under 7 days it runs autonomously, because a stale cert that expires at 3 a.m. is worse than a controlled restart at 3 p.m. Above 7 days it posts for approval, with the two serials side by side.
  • retrigger_renewal is once per certificate per 24 hours, tracked in the agent's own state. This is the rate-limit protection the prompt cannot guarantee.
  • All mutations sit behind the circuit breaker. A certificate agent that restarts four deployments during an unrelated incident has made the incident worse, so the SLO burn gate applies here like everywhere else.

The restart itself is the same patch kubectl rollout restart sends, applied only when the executor has re-verified the serial mismatch a second time immediately before acting:

def rollout_restart(ns, kind, name, expected_secret_serial, host):
    if served_serial(host) == expected_secret_serial:
        return "skipped: served cert already current"
    body = {"spec": {"template": {"metadata": {"annotations": {
        "certagent.devtocash.com/restartedAt": datetime.now(timezone.utc).isoformat()}}}}}
    getattr(apps, f"patch_namespaced_{kind}")(name, ns, body)

A Real Run

A 40-namespace EKS cluster, 61 TLS Secrets, 38 of them managed by cert-manager. The inventory produced a shortlist of six:

shop/shop-tls          managed   days_left=19  ready=False  acme_http01_unreachable      -> propose_ingress_fix_pr
api/api-tls            managed   days_left=27  ready=True   renewed_not_reloaded          -> rollout_restart (approval)
legacy/portal-tls      unmanaged days_left=9   ready=n/a    unmanaged_secret_expiring     -> ask_owner
admin/grafana-tls      managed   days_left=3   ready=False  acme_rate_limited             -> wait_for_rate_limit
edge/www-tls           managed   days_left=41  ready=True   served_cert_is_external       -> report_only
ml/jupyter-tls         managed   days_left=6   ready=True   renewed_not_reloaded          -> rollout_restart (auto)

The shop failure was a new path-based Ingress rule that shadowed the ACME challenge path with a 404 from the storefront; the PR the agent opened added the /.well-known/acme-challenge/ route back above it. The grafana cert had been recreated by a Helm upgrade six times that week, one Order per upgrade, and had hit the duplicate-certificate limit; the honest answer was to wait two days, and the agent said so with the date. The legacy portal turned out to be a Secret pasted in by a contractor in 2024 with a 2-year lifetime, which nobody knew until the agent asked. The www cert lives on the CloudFront distribution in front of the cluster and was never in scope. Two restarts, one PR, one Slack message, one wait, and nothing expired.

Honest Limits

The served-serial check only works for hosts the agent can reach on 443 from inside the cluster, and only for hosts that resolve to the same place users hit. Split-horizon DNS or an internal-only Ingress class silently drops those hosts from class-four detection, and you should log the count of unreachable hosts as its own metric.

The unmanaged-Secret scan tells you a cert is expiring, not who owns it. Without a service catalog to resolve the namespace to a team, ask_owner degrades into a message in a shared channel, which is where certificate expiries have always gone to die.

And the fifth class stays mostly out of reach. Webhook caBundles, kubelet serving certificates, and service-mesh CA roots each expire on their own schedules, and the upgrade readiness agent is the better place for the webhook check because that is when it bites. This agent earns its keep on the boring majority: the certificates cert-manager was supposed to renew, and quietly did not.

#ai#sre#kubernetes#reliability#automation#security
D
DevToCashAuthor

Senior DevOps/SRE Engineer · 10+ years · Professional Trader (IDX, Crypto, US Equities)

I write about real infrastructure patterns and trading strategies I use in production and in live markets. No courses, no affiliate hype — just documentation of what actually works.

More about me →