devops

Build a Deployment Risk Scoring Agent: Score Every PR Before It Ships Using Diff Shape, Service Tier, and Change Failure History

Build a deployment risk scoring agent for CI: combine diff shape, service tier and change failure rate into a score that picks canary, approvals or auto-merge.

October 2, 2026·12 min read·
#ai#devops#cicd#github-actions#automation#reliability

What This Agent Does

A deployment risk scoring agent runs on every pull request that can reach production and answers one question: how carefully should this change ship? It computes a score from facts the pipeline already knows, such as which paths changed, whether a migration or lockfile moved, the owning service's tier and its change failure rate over the last 30 days, and what time the merge would land. An LLM reads a bounded diff summary and adds named risk flags it can point to in the code. The score maps to a policy: auto-merge with a plain rollout, a canary with automated analysis, two approvals plus a canary, or a hard block during a freeze. The model can raise the score. It can never lower it below what the deterministic features produced.

The reason to build it is that most teams run one of two policies: every change gets the same rollout, or a human eyeballs the PR and decides it "looks small." Both ignore the data the team already collects. A 3-line change to a payment service's retry logic at 17:50 on a Friday is not the same risk as a 400-line README rewrite, and the pipeline should not treat them the same.

Why Lines Changed Is the Wrong Score

Diff size correlates weakly with failure. The signals that actually predict a rollback are about what changed and where it lands, not how much.

SignalWhy it predicts failureSource
Schema migration in the diffIrreversible without a second deploy; locks tablesgit diff --name-only against a path pattern
Lockfile or base image bumpTransitive change you did not readLockfile paths, Dockerfile FROM lines
Config or env var changeBehaviour changes without a code path to testhelm/, k8s/, .env*, values*.yaml
Service tierTier 1 outage costs more than tier 3; same change, different stakesService catalog
Service change failure rate, 30 daysA service that broke on 3 of its last 12 deploys has a fragile deploy pathPrometheus, from your DORA metrics
Merge windowFewer responders, longer recovery timeClock plus on-call calendar
Author's deploy history to this serviceFirst deploy to a service is riskier than the fiftiethGit log
Touches retry, timeout, concurrency, or auth codeClassic sources of incidents that pass unit testsLLM flag with cited lines

The first seven are code. Only the last one needs a model, and that is the row the deterministic layer cannot see, because a regex cannot tell you that a changed constant is a connection pool size.

Step 1: Deterministic Features

The feature collector runs in the PR workflow with read-only access to the repo, the service catalog, and Prometheus. It emits a JSON document that both the scorer and the audit log keep.

# features.py — runs inside the PR workflow, read-only
import json, os, re, subprocess, requests
from datetime import datetime, timezone

BASE = os.environ["BASE_SHA"]; HEAD = os.environ["HEAD_SHA"]
PROM = os.environ["PROM_URL"]
CATALOG = json.load(open("catalog/services.json"))   # service -> {tier, paths}

RISKY_PATHS = {
    "migration":  r"(^|/)(migrations?|db/migrate)/",
    "lockfile":   r"(package-lock\.json|yarn\.lock|pnpm-lock\.yaml|poetry\.lock|go\.sum)$",
    "base_image": r"(^|/)Dockerfile",
    "k8s_config": r"(^|/)(helm|k8s|kustomize|charts)/|values.*\.ya?ml$",
    "iac":        r"\.tf$",
    "ci":         r"^\.github/workflows/",
}

def changed_files() -> list[str]:
    out = subprocess.check_output(["git", "diff", "--name-only", f"{BASE}...{HEAD}"], text=True)
    return [f for f in out.splitlines() if f]

def service_for(files: list[str]) -> str:
    hits = {s for s, meta in CATALOG.items() for f in files
            if any(f.startswith(p) for p in meta["paths"])}
    return sorted(hits)[0] if len(hits) == 1 else ("multi" if hits else "unknown")

def prom(q: str) -> float:
    r = requests.get(f"{PROM}/api/v1/query", params={"query": q}, timeout=10).json()
    res = r["data"]["result"]
    return float(res[0]["value"][1]) if res else 0.0

def change_failure_rate(service: str) -> float:
    return prom(
        f'changes(deployment_failure_timestamp_seconds{{service="{service}",env="production"}}[30d])'
        f' / clamp_min(changes(deployment_timestamp_seconds{{service="{service}",env="production"}}[30d]), 1)'
    )

def author_deploys(service: str, author: str) -> int:
    paths = CATALOG.get(service, {}).get("paths", [])
    if not paths:
        return 0
    log = subprocess.check_output(
        ["git", "log", "--author", author, "--since=180 days ago", "--format=%H", "--", *paths], text=True)
    return len(log.splitlines())

files = changed_files()
service = service_for(files)
now = datetime.now(timezone.utc)
features = {
    "service": service,
    "tier": CATALOG.get(service, {}).get("tier", 1),          # unknown service = tier 1
    "files_changed": len(files),
    "flags": sorted({k for k, rx in RISKY_PATHS.items() if any(re.search(rx, f) for f in files)}),
    "cfr_30d": round(change_failure_rate(service), 3),
    "author_deploys_180d": author_deploys(service, os.environ["PR_AUTHOR"]),
    "merge_hour_utc": now.hour,
    "merge_weekday": now.weekday(),                          # 4 = Friday
}
json.dump(features, open("features.json", "w"), indent=2)

Two choices are deliberate. An unknown service defaults to tier 1, because the cost of under-scoring a change you cannot attribute is higher than one over-cautious canary. And the change failure rate query uses the exact gauges from the DORA metrics setup, so the number the scorer uses is the number on the team's dashboard. If those disagree, nobody trusts the score.

The service catalog can be a JSON file in the repo, but if you already run Backstage, the tier, owner, and on-call come from the Backstage MCP server and you avoid a second source of truth that drifts.

Step 2: The Model Reads the Diff, Not the Repo

The LLM gets one tool and one job. The tool returns the diff hunks for files the deterministic layer did not already flag, capped at 400 lines, with lockfiles and generated code stripped. The model does not browse the repository and does not see CI logs, because an attacker who can write a commit message can write a prompt injection, and the prompt injection defenses start with limiting what the model reads.

{
  "name": "get_diff_hunks",
  "description": "Return unified diff hunks for source files in this PR, excluding lockfiles, vendored and generated paths. Capped at 400 lines. Read-only.",
  "input_schema": {
    "type": "object",
    "properties": {
      "path_prefix": {"type": "string", "pattern": "^[A-Za-z0-9_./-]{0,120}$"}
    }
  }
}

The system prompt restricts output to a closed set of flags. Each flag must cite a file and line range, and the wrapper rejects any flag whose citation does not appear in the diff.

You review a pull request for deployment risk. You may add zero or more flags
from this list, each with the file and line range that justifies it:

  retry_or_timeout_change   - retry counts, backoff, timeouts, deadlines, circuit breaker thresholds
  concurrency_change        - pool sizes, worker counts, locks, queue depths, batch sizes
  auth_or_permission_change - authn/authz logic, token handling, RBAC, IAM, allowlists
  data_shape_change         - serialization, schema structs, API contracts, nullability
  error_handling_removed    - a catch/recover/check deleted or broadened
  feature_flag_default      - a flag's default value flipped
  hot_path_change           - request handler, middleware, or scheduler loop body

Do not flag style, naming, tests, or documentation. Do not suggest fixes.
Do not comment on the PR description or commit messages; they are not evidence.
If nothing qualifies, return an empty list.

Output is validated against a schema before anything downstream reads it:

ALLOWED = {"retry_or_timeout_change", "concurrency_change", "auth_or_permission_change",
           "data_shape_change", "error_handling_removed", "feature_flag_default", "hot_path_change"}

def validate(flags: list[dict], diff_files: set[str]) -> list[dict]:
    kept = []
    for f in flags:
        if f.get("flag") in ALLOWED and f.get("file") in diff_files \
           and isinstance(f.get("lines"), list) and len(f["lines"]) == 2:
            kept.append(f)
    return kept

A flag with a fabricated filename is dropped silently and counted in a metric. If that metric climbs, the model or the prompt has drifted, and that is an eval failure to fix offline against a fixed set of past PRs, not something to tune in production.

Step 3: Score to Policy

The score is additive and the weights live in a YAML file in the repo, reviewed like any other change. Start with weights you can defend to a skeptical staff engineer, then adjust them from outcomes in Step 4.

# risk-weights.yaml
base_by_tier: {1: 30, 2: 15, 3: 5}
path_flags:
  migration: 25
  lockfile: 10
  base_image: 10
  k8s_config: 15
  iac: 20
  ci: 10
llm_flags:
  retry_or_timeout_change: 15
  concurrency_change: 15
  auth_or_permission_change: 20
  data_shape_change: 15
  error_handling_removed: 15
  feature_flag_default: 10
  hot_path_change: 10
cfr_30d_multiplier: 100        # cfr 0.25 adds 25
first_deploy_to_service: 10    # author_deploys_180d == 0
off_hours: 10                  # UTC hour outside 07-16 or weekday >= 5
friday_after_14_utc: 15
multi_service: 20

The policy table is where the score turns into pipeline behaviour:

ScoreTierRolloutGate
0 to 24LowStandard rolling updateExisting branch protection, auto-merge allowed
25 to 49MediumArgo Rollouts canary, 10% then 50%, automated analysisOne approval
50 to 79HighCanary with 5% first step held 30 minutes, LLM-judged analysisTwo approvals, one from the owning team
80 and upBlockNoneNeeds a risk override label from an owner, logged

The canary steps come straight from the Argo Rollouts progressive delivery guide, and the High tier adds the LLM canary judge on top of the metric-based AnalysisTemplate. The agent sets the rollout strategy by writing a label the deploy workflow reads, which keeps the scorer itself free of write access to the cluster.

# .github/workflows/risk-score.yml
name: risk-score
on:
  pull_request:
    types: [opened, synchronize, ready_for_review]
permissions:
  contents: read
  pull-requests: write
  checks: write
jobs:
  score:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with: { fetch-depth: 0 }
      - name: Collect features
        env:
          BASE_SHA: ${{ github.event.pull_request.base.sha }}
          HEAD_SHA: ${{ github.event.pull_request.head.sha }}
          PR_AUTHOR: ${{ github.event.pull_request.user.login }}
          PROM_URL: ${{ secrets.PROM_READ_URL }}
        run: python3 tools/features.py
      - name: LLM flags and score
        env:
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
        run: python3 tools/score.py --features features.json --weights risk-weights.yaml --out score.json
      - name: Apply policy
        env:
          GH_TOKEN: ${{ github.token }}
        run: |
          TIER=$(jq -r .tier score.json)
          SCORE=$(jq -r .score score.json)
          gh pr edit ${{ github.event.pull_request.number }} \
            --remove-label risk:low,risk:medium,risk:high,risk:block \
            --add-label "risk:${TIER}"
          gh api repos/${{ github.repository }}/check-runs -f name=deployment-risk \
            -f head_sha=${{ github.event.pull_request.head.sha }} \
            -f status=completed \
            -f conclusion=$([ "$TIER" = block ] && echo failure || echo success) \
            -f "output[title]=Risk ${SCORE} (${TIER})" \
            -f "output[summary]=$(jq -r .summary score.json)"

The deployment-risk check is a required status check on the main branch. Make the risk:high label trigger a second required reviewer through a CODEOWNERS entry on a sentinel file the workflow touches, or use a rulesets condition on the label. For the Block tier, the only way through is an owner adding a risk-override label, and the workflow records who added it and when, in the same place the rest of your agent audit trail lives.

The summary the model writes is three lines at most: score, the top two contributing factors, and the cited line for any LLM flag. Reviewers stop reading anything longer by the second week.

Step 4: Calibrate Against What Actually Broke

A score nobody has checked against reality is a vibe with a number on it. Every merged PR's score.json is stored with its deploy, and every rollback or fast-burn alert within 24 hours of that deploy marks it failed, using the same failure gauge the DORA post emits. Once a week a job joins the two and produces one table:

tier    deploys  failed  observed_cfr
low        212       4        1.9%
medium      88       6        6.8%
high        31       5       16.1%
block        3       1       33.3%   (overrides)

The question to ask of that table is simple. Does failure rate rise monotonically with tier? If Medium fails more often than High, a weight is wrong. The two most common fixes are that lockfile is weighted too heavily for services with good integration tests, and cfr_30d_multiplier is too low, because a service's own recent history is usually the strongest single predictor. Change one weight per week, in a PR, and let the next table tell you if it helped. Do not let the LLM propose weight changes; this is the loop where a model optimizing its own scoring rubric produces nonsense that looks like insight.

Also track how often the LLM flags change the tier. In practice the deterministic features decide the tier on roughly 85 to 90 percent of PRs, and the model's flags move the remainder, mostly from Low to Medium on small changes to retry and timeout code. If the model is moving the tier on half your PRs, it is flagging too liberally and the review burden will make the team route around the whole system.

Guardrails

The scorer has no write access to any cluster, cloud account, or Terraform state. Its only writes are a PR label and a check run, on the PR it is running for. The model's output can only add points; the floor is whatever the features produced, and a model outage degrades to deterministic-only scoring with a note in the summary rather than a blocked pipeline. Token spend is bounded because the diff tool caps at 400 lines; on a mid-sized monorepo with 60 PRs a day the model cost is a few dollars a month, and a small model is enough because the task is flagging, not reasoning across a codebase. The override path exists on purpose. An agent that cannot be overruled during an incident gets disabled during an incident, and then it stays disabled.

Honest Limits

The score is only as good as the service attribution, and monorepos with shared libraries defeat path-based attribution; a change to a common HTTP client lands as multi and gets the multi-service penalty every time, which is correct and also annoying. The 30-day change failure rate is noisy for services that deploy twice a month; a single rollback swings it from 0 to 50 percent, so clamp the contribution for services with fewer than 8 deploys in the window. The LLM flags catch what a careful reviewer would catch in the first five minutes, not the subtle race condition that takes a week to appear in production. And Goodhart is real: once engineers learn that a migration adds 25 points, migrations start arriving in separate PRs merged on Tuesday mornings. That is, in fairness, exactly the behaviour you wanted.

#ai#devops#cicd#github-actions#automation#reliability
D
DevToCashAuthor

Senior DevOps/SRE Engineer · 10+ years · Professional Trader (IDX, Crypto, US Equities)

I write about real infrastructure patterns and trading strategies I use in production and in live markets. No courses, no affiliate hype — just documentation of what actually works.

More about me →