What This Agent Does
A Kubernetes manifest review agent runs on every pull request that changes a chart, an overlay, or a raw manifest, and answers one question per workload: will this object behave in the cluster the way the author thinks it will. It renders the Helm chart or Kustomize overlay the same way the deploy job does, so it reviews what the API server will actually receive. Three deterministic linters then do the finding: kubeconform for schema validity, kube-score for runtime practices, and pluto for APIs that stop existing on the next upgrade. A read-only enricher adds what the linters cannot know, such as the namespace LimitRange, the existing Deployment's current requests, and the Vertical Pod Autoscaler recommendation if one exists. The LLM receives only those enriched findings and writes the review: what breaks, when, and the exact YAML patch. It never applies anything, never invents a memory number, and blocks the merge only on findings that admission would reject anyway.
The deployment risk scoring agent counts "a manifest changed" as a risk flag. This agent reads the manifest.
Why Admission Control Is Too Late
If you already enforce Kyverno policies in Enforce mode, a Deployment without resource limits never reaches the cluster. That is correct and you should keep it. The problem is where the author finds out. Argo CD syncs, the API server returns a 400 with a policy message, the sync goes Degraded, and the author learns about it an hour after merge from a Slack alert that quotes a Kyverno rule name. Then they open a second PR.
Moving the same checks into the PR turns a post-merge outage-shaped event into a review comment with a patch. It also catches the class of problems admission does not reject, because they are legal YAML with bad consequences:
| Finding | Legal? | What actually happens |
|---|---|---|
No readinessProbe | Yes | Traffic hits the pod during startup; rolling updates ship 502s |
livenessProbe identical to readiness, 1s timeout | Yes | Slow GC pause becomes a restart loop |
No memory limit | Yes | One leak evicts neighbours; see the OOMKilled guide for the other failure |
Memory limit below the VPA lower bound | Yes | OOMKilled on the first real request |
image: app:latest | Yes | Rollback restores the same tag and therefore the same bug |
replicas: 2 with no PodDisruptionBudget | Yes | A node drain takes both replicas |
apiVersion: policy/v1beta1 PDB | Yes, today | Fails the day the cluster upgrades |
Every row has a known fix. The agent's job is to attach the fix to the line that needs it, with the numbers taken from the cluster rather than from the model.
Step 1: Render What the Cluster Will See
Linting the chart source is pointless. Values files change defaults, Kustomize patches add or remove probes, and a templated {{ .Values.resources }} is empty in half the overlays. The job renders every environment the PR affects with the same commands the deploy pipeline uses.
# .github/workflows/manifest-review.yml (excerpt)
on:
pull_request:
paths: ["charts/**", "deploy/**", "k8s/**"]
jobs:
review:
runs-on: ubuntu-latest
permissions:
contents: read
pull-requests: write
id-token: write # OIDC to the cluster for the read-only enricher
steps:
- uses: actions/checkout@v4
- name: Render
run: |
mkdir -p rendered
for env in staging prod; do
helm template api charts/api \
-f charts/api/values.yaml -f charts/api/values-${env}.yaml \
--namespace api --kube-version 1.31.0 > rendered/${env}.yaml
done
kustomize build deploy/overlays/prod >> rendered/prod.yaml
The --kube-version flag matters. Helm templates that branch on .Capabilities.KubeVersion will otherwise render against Helm's built-in default and produce objects the real cluster has never seen.
Step 2: Three Linters, One Findings File
Each tool covers a different failure class and all three emit JSON, so the merge step is a filter over a list rather than a pile of regexes against terminal output.
# Schema: does this object exist in the API at this version, with these fields?
kubeconform -kubernetes-version 1.31.0 -strict -summary -output json \
-schema-location default \
-schema-location 'https://raw.githubusercontent.com/datreeio/CRDs-catalog/main/{{.Group}}/{{.ResourceKind}}_{{.ResourceAPIVersion}}.json' \
rendered/*.yaml > lint/schema.json || true
# Practices: probes, limits, tags, PDBs, security context
kube-score score rendered/*.yaml --output-format json \
--ignore-test container-image-pull-policy > lint/score.json || true
# Deprecations: anything removed by the target version
pluto detect-files -d rendered -o json --target-versions k8s=v1.32.0 > lint/pluto.json || true
Three details that save a week of false positives. The second -schema-location points kubeconform at the community CRD catalog, so an Argo Rollout or a ServiceMonitor validates instead of being skipped. The kube-score --ignore-test drops the pull-policy check, because with immutable digests that check is noise. And --target-versions for pluto is set one minor version ahead of the cluster, so the PR that would break the next upgrade fails now, which is the whole point of the upgrade readiness agent applied one object at a time.
The normaliser flattens all three outputs into one shape:
{
"env": "prod",
"kind": "Deployment",
"name": "api",
"namespace": "api",
"container": "api",
"check": "container-resources",
"severity": "critical",
"message": "Container has no memory limit",
"source": "kube-score"
}
Severity is assigned by the normaliser from a fixed table, not by the model. Schema errors and anything the production Kyverno policy set would reject are critical. Missing probes, missing PDB on a multi-replica Deployment, and mutable tags are high. Everything else is medium.
Step 3: Enrich With Numbers From the Cluster
A finding that says "no memory limit" is a lint. A finding that says "no memory limit; the current Deployment runs at 310 MiB p95 and the VPA lower bound is 384 MiB" is a review. The enricher connects with a ServiceAccount that can get and list a handful of resource types and nothing else, following the same posture as the least-privilege RBAC setup for every agent on this site.
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: manifest-reviewer
rules:
- apiGroups: ["", "apps", "policy"]
resources: ["limitranges", "resourcequotas", "deployments", "statefulsets", "poddisruptionbudgets"]
verbs: ["get", "list"]
- apiGroups: ["autoscaling.k8s.io"]
resources: ["verticalpodautoscalers"]
verbs: ["get", "list"]
- apiGroups: ["metrics.k8s.io"]
resources: ["pods"]
verbs: ["list"]
# enrich.py — read-only; KUBECONFIG comes from the OIDC step
import json, subprocess
def kget(*args):
out = subprocess.run(["kubectl", "get", *args, "-o", "json"],
capture_output=True, text=True, timeout=20)
return json.loads(out.stdout) if out.returncode == 0 else None
def enrich(f: dict) -> dict:
ns, name = f["namespace"], f["name"]
live = kget(f["kind"].lower(), name, "-n", ns)
if live:
for c in live["spec"]["template"]["spec"]["containers"]:
if c["name"] == f.get("container"):
f["live_resources"] = c.get("resources", {})
f["live_replicas"] = live["spec"].get("replicas")
vpa = kget("vpa", "-n", ns)
for v in (vpa or {}).get("items", []):
if v["spec"]["targetRef"]["name"] == name:
for rec in v.get("status", {}).get("recommendation", {}) \
.get("containerRecommendations", []):
if rec["containerName"] == f.get("container"):
f["vpa"] = {k: rec[k] for k in ("lowerBound", "target", "upperBound") if k in rec}
lr = kget("limitrange", "-n", ns)
f["limitrange_defaults"] = [
item.get("default")
for obj in (lr or {}).get("items", [])
for item in obj["spec"].get("limits", [])
if item.get("type") == "Container"
]
pdbs = kget("pdb", "-n", ns)
f["pdb_exists"] = any(p["spec"].get("selector") for p in (pdbs or {}).get("items", []))
return f
Two outcomes come straight from this data without the model. If the namespace has a LimitRange with a default memory limit, the "no memory limit" finding is downgraded to medium, because the pod will get one at admission. And if a VPA exists, its target becomes the number the patch proposes. A model asked to pick a memory limit will pick a round number. A VPA picked it from the last eight days of usage.
Step 4: The Model Writes the Review, Not the Decision
The LLM sees the enriched findings and the relevant slice of rendered YAML, and it is forced into a structured tool call so the output is always parseable. The system prompt pins the one rule that keeps this safe: every number in a patch must appear in the input.
# review.py
import json, anthropic
SYSTEM = """You review rendered Kubernetes manifests before merge.
For each finding, explain what happens at runtime, when it will be noticed, and give a
YAML patch the author can paste. Rules:
- Use only resource values present in the finding (live_resources, vpa, limitrange_defaults).
If none are present, say so and give the patch with a TODO placeholder, never a guess.
- For probes, use the container's declared port and the /healthz or /readyz path only if it
appears in the rendered manifest; otherwise ask for the path.
- Do not change the severity you are given. Do not comment on findings you were not given."""
def review(findings: list[dict], yaml_slice: str) -> dict:
client = anthropic.Anthropic()
msg = client.messages.create(
model="claude-sonnet-5", max_tokens=4000, system=SYSTEM,
tools=[{
"name": "manifest_review",
"description": "Structured review of one rendered workload",
"input_schema": {
"type": "object",
"properties": {
"findings": {"type": "array", "items": {
"type": "object",
"properties": {
"check": {"type": "string"},
"runtime_effect": {"type": "string"},
"noticed_when": {"type": "string"},
"patch_yaml": {"type": "string"}},
"required": ["check", "runtime_effect", "noticed_when", "patch_yaml"]}},
"summary": {"type": "string"}},
"required": ["findings", "summary"]}}],
tool_choice={"type": "tool", "name": "manifest_review"},
messages=[{"role": "user", "content":
"Rendered manifest (YAML):\n" + yaml_slice +
"\n\nFindings (JSON):\n" + json.dumps(findings, indent=1)}])
return next(b.input for b in msg.content if b.type == "tool_use")
A representative comment the harness posts, with the patch taken from the VPA target in the finding:
critical · Deployment/api (prod) · container-resources No memory limit. The live Deployment requests 256Mi and the VPA target is 412Mi with an upper bound of 640Mi. Without a limit this pod is the first candidate for eviction under node pressure and it inherits no default from a LimitRange (none in namespace
api). Noticed: the next time a node fills.resources: requests: cpu: 200m memory: 412Mi limits: memory: 640Mi
Note what the patch does not do. It does not set a CPU limit, because kube-score flags CPU limits as optional and throttling is usually worse than the problem it prevents. The prompt did not have to say this; the finding simply never asked for one.
Step 5: The Merge Gate
# gate.py
import json, sys
findings = json.load(open("lint/enriched.json"))
labels = set(json.load(open("pr.json"))["labels"])
crit = [f for f in findings if f["severity"] == "critical" and f["env"] == "prod"]
if crit and "manifest-risk-accepted" not in labels:
print(f"{len(crit)} critical prod findings; add label manifest-risk-accepted to override")
sys.exit(1)
Only critical findings in the prod render block. A missing probe in staging is a comment. A schema error in prod is a failed check, and it would have failed at sync time anyway, so nobody is being asked to accept a stricter rule than the cluster already has. The override label is logged by GitHub with the name of whoever applied it, which is the audit trail.
To keep the gate honest with your actual admission policy, run the Kyverno CLI against the render and feed its failures into the normaliser as critical:
kyverno apply policies/prod/ --resource rendered/prod.yaml --policy-report > lint/kyverno.yaml
Now the question "would admission reject this" is answered by the same policy files admission uses, not by a re-implementation of them in kube-score terms.
What Goes Wrong
Rendering with the wrong values. If CI renders with values.yaml only and prod deploys with a secrets overlay, the review is for an object that never ships. Render every overlay the deploy job renders, and fail the job if the deploy job's values list and the review job's list drift apart. A one-line diff of the two helm template command strings in a check step is enough.
kube-score on Jobs and CronJobs. A one-shot Job legitimately has no readiness probe and no PDB. Ignore pod-probes and deployment-has-poddisruptionbudget for kind: Job and kind: CronJob in the normaliser, or the CronJob health agent will be competing with this one for the author's patience.
Stale VPA recommendations. A VPA in Off mode on a workload that was rewritten last month recommends memory for the old code. Include the recommendation's age from status.conditions in the finding, and have the prompt flag any recommendation older than 14 days as "verify before applying" instead of pasting it into the patch.
Model drift on patches. Once a quarter, replay the last 50 reviews through the current model and diff the patch_yaml fields against what was merged. If the model starts proposing values that were not in its input, the prompt rule has stopped working and the enricher output needs to be the only place numbers can come from, enforced by a regex over the patch before posting.
Token cost. A rendered prod manifest for a mid-size chart is 2,000 to 6,000 tokens; the findings are a few hundred. At roughly one review per PR with Sonnet-class pricing, this is cents per PR. The expensive part is the read-only cluster calls on a cold runner, which is why the enricher batches by namespace rather than calling kubectl per finding.
Where This Fits
This is the third review agent in a pattern that started with Terraform plans and continued with schema migrations. Each one has the same shape: a deterministic detector, a read-only enricher that pulls live numbers, a model that writes the explanation and the patch from those numbers only, and a gate that blocks on the subset of findings the next system in line would have rejected anyway. The model is the cheapest part and the least trusted. That is deliberate. The probes guide at liveness, readiness, and startup probes covers what the patches in this agent are trying to get right.