devops

Build a GitHub Actions Cost Agent: Find Wasted Runner Minutes, Cache Misses, and Superseded Runs Before the Invoice Does

Build a GitHub Actions cost agent that finds wasted runner minutes: superseded runs, cache misses, rounding tax, and oversized runners, then opens the fix PR.

October 4, 2026·13 min read·
#ai#devops#github-actions#cicd#finops#cost-optimization#automation

What This Agent Does

A GitHub Actions cost agent reads your organization's Actions bill per repository and SKU, reconstructs where every billed minute came from at the job level, and classifies the waste into six patterns that each have a known one-file fix: superseded runs that should have been cancelled, caches that never hit, jobs billed for a full minute to do twenty seconds of work, runners with more cores than the job can use, macOS or Windows jobs that would run fine on Linux, and jobs with no timeout burning the six-hour default. Deterministic code finds and sizes the waste. The model does one thing with it: write the workflow patch and the PR description, which actionlint checks before a human reviews it. The agent never edits a workflow directly and never deletes a cache.

Actions spend is the CI bill that nobody owns. It is not in the cloud account, so the FinOps review skips it. It is per-minute, so no single job looks expensive. And the organization billing page shows minutes by repository but never by cause, which is exactly the gap this agent fills.

Where the Minutes Actually Go

GitHub bills hosted runners per job, rounded up to the next whole minute, with a per-platform price. Public repositories on standard runners are free. Private repositories draw from the plan's included minutes first and are billed after that. The multipliers are what make the patterns below expensive.

RunnerPrice per minuteMultiplier vs Linux 2-core
Linux 2-core$0.0081x
Linux 4-core$0.0162x
Linux 8-core$0.0324x
Linux 16-core$0.0648x
Windows 2-core$0.0162x
macOS 3-core$0.0810x

The six waste patterns, with the API fingerprint the agent uses to detect each one:

PatternHow the agent detects itFix
Superseded runTwo runs of the same workflow on the same branch overlap in time and the older one was not cancelledconcurrency with cancel-in-progress
Rounding taxMany jobs with a wall clock under 60 seconds, each billed as 1 minuteMerge short jobs, shrink the matrix
Cache missJob log contains Cache not found for input keys more often than Cache restored from keyKey on a lockfile hash, add restore-keys
Oversized runnerLarger-runner label on a job whose median duration is dominated by checkout and installDownsize and benchmark
Platform multiplierLint, docs, or unit-test jobs on macos-* or windows-* labelsMove to ubuntu-latest, keep one real macOS job
No timeoutA job that normally takes minutes with an outlier near 360 minutestimeout-minutes

Nothing in that table needs a model. The model earns its place later, when a finding has to become a correct YAML change in a workflow file the agent has never seen.

Step 1: Pull the Bill by Repository and SKU

The enhanced billing usage endpoint returns one line per day, repository, and SKU, with the net amount after included minutes are applied. That is the number finance sees, so start there and rank repositories by it.

# bill.py — read-only; org-level token with billing read access
import os, collections, requests
from datetime import date

ORG = os.environ["GH_ORG"]
H = {"Authorization": f"Bearer {os.environ['GH_TOKEN']}",
     "Accept": "application/vnd.github+json",
     "X-GitHub-Api-Version": "2022-11-28"}

def actions_minutes(year: int, month: int) -> list[dict]:
    r = requests.get(f"https://api.github.com/organizations/{ORG}/settings/billing/usage",
                     headers=H, params={"year": year, "month": month}, timeout=30)
    r.raise_for_status()
    return [u for u in r.json()["usageItems"]
            if u["product"] == "Actions" and u["unitType"].lower() == "minutes"]

today = date.today()
by_repo_sku = collections.defaultdict(lambda: {"minutes": 0.0, "net_usd": 0.0})
for u in actions_minutes(today.year, today.month):
    k = (u["repositoryName"], u["sku"])
    by_repo_sku[k]["minutes"] += u["quantity"]
    by_repo_sku[k]["net_usd"] += u["netAmount"]

for (repo, sku), v in sorted(by_repo_sku.items(), key=lambda kv: -kv[1]["minutes"])[:15]:
    print(f"{v['net_usd']:8.2f} USD {v['minutes']:9.0f} min  {repo:40s} {sku}")

Rank by minutes, not dollars, when the organization is still inside its included allowance. A repository burning 2,800 of a 3,000-minute allowance costs nothing on this month's invoice and everything on next month's, and the queue time it inflicts on every other repository is a cost the bill never shows.

The SKU column is the first finding on its own. A repository whose top SKU is Actions macOS for a service that ships a Linux container has a ten-times multiplier on jobs that almost certainly do not need it.

Step 2: Reconstruct Billable Time per Job

The billing API stops at the repository. Everything below that comes from the workflow runs and jobs endpoints. Per job you get the runner labels and the start and completion timestamps, which is enough to compute what GitHub billed, because the rounding rule is public.

# jobs.py — per-job billable reconstruction for one repository, last 14 days
import math, collections, requests
from datetime import datetime, timedelta, timezone

def gh(url: str, **params):
    r = requests.get(url, headers=H, params=params, timeout=30)
    r.raise_for_status()
    return r.json()

def ts(s: str) -> datetime:
    return datetime.fromisoformat(s.replace("Z", "+00:00"))

def runs(repo: str, days: int = 14) -> list[dict]:
    since = (datetime.now(timezone.utc) - timedelta(days=days)).strftime("%Y-%m-%d")
    page, out = 1, []
    while True:
        batch = gh(f"https://api.github.com/repos/{repo}/actions/runs",
                   created=f">={since}", per_page=100, page=page)["workflow_runs"]
        out += batch
        if len(batch) < 100:
            return out
        page += 1

def job_rows(repo: str, run: dict) -> list[dict]:
    rows = []
    jobs = gh(f"https://api.github.com/repos/{repo}/actions/runs/{run['id']}/jobs",
              per_page=100)["jobs"]
    for j in jobs:
        if not (j["started_at"] and j["completed_at"]):
            continue
        secs = (ts(j["completed_at"]) - ts(j["started_at"])).total_seconds()
        rows.append({
            "run_id": run["id"], "job_id": j["id"], "workflow": run["name"],
            "path": run["path"], "event": run["event"], "branch": run["head_branch"],
            "job": j["name"], "labels": ",".join(j["labels"]),
            "secs": secs, "billed_min": math.ceil(secs / 60),
        })
    return rows

def rounding_tax(rows: list[dict]) -> list[tuple]:
    tax = collections.Counter()
    for r in rows:
        tax[(r["path"], r["job"])] += r["billed_min"] - r["secs"] / 60
    return tax.most_common(10)

def superseded(all_runs: list[dict]) -> list[dict]:
    by_key = collections.defaultdict(list)
    for r in all_runs:
        if r["event"] in ("push", "pull_request") and r["conclusion"] != "cancelled":
            by_key[(r["path"], r["head_branch"])].append(r)
    found = []
    for (path, branch), rs in by_key.items():
        rs.sort(key=lambda r: r["run_started_at"])
        for older, newer in zip(rs, rs[1:]):
            if ts(newer["run_started_at"]) < ts(older["updated_at"]):
                found.append({"path": path, "branch": branch,
                              "wasted_run": older["id"], "replaced_by": newer["id"]})
    return found

def no_timeout_suspects(rows: list[dict]) -> list[tuple]:
    by_job = collections.defaultdict(list)
    for r in rows:
        by_job[(r["path"], r["job"])].append(r["secs"])
    out = []
    for key, secs in by_job.items():
        secs.sort()
        median = secs[len(secs) // 2]
        if secs[-1] > 3600 and secs[-1] > 6 * max(median, 1):
            out.append((key, median, secs[-1]))
    return out

Three details in that code are deliberate. The superseded check keys on the workflow path rather than its display name, because two workflows can share a name and the path is what concurrency groups would key on. It excludes runs that were already cancelled, so a repository that has the fix gets no finding. And the timeout heuristic looks for an outlier at least six times the median and over an hour, because a job that is slow every time is a performance problem, not a missing timeout.

The rounding tax is the number that surprises teams. A matrix of twelve lint jobs that each take 25 seconds is billed as 12 minutes for 5 minutes of work. On a 4-core runner that is $0.19 per push, and on a repository with 60 pushes a day it is about $340 a month for a matrix that could be one job with a loop. The agent reports the tax as billed minutes minus wall-clock minutes per job, summed across the window, so the PR can cite it.

Step 3: Read Cache Outcomes from the Job Logs

The cache API tells you what is stored, not what was used. Whether a restore hit or missed is only in the job log, and the two lines actions/cache prints are stable enough to grep.

# cache.py — classify cache outcomes from job logs
import re, requests

MISS = re.compile(r"Cache not found for input keys: (.+)")
HIT = re.compile(r"Cache restored from key: (.+)")
CHURN = re.compile(r"[0-9a-f]{40}|run_id|run_number|github\.sha")

def cache_outcome(repo: str, job_id: int) -> dict:
    r = requests.get(f"https://api.github.com/repos/{repo}/actions/jobs/{job_id}/logs",
                     headers=H, timeout=60, allow_redirects=True)
    text = r.text if r.ok else ""
    misses = MISS.findall(text)
    return {
        "hits": len(HIT.findall(text)),
        "misses": len(misses),
        "churn_keys": [k for k in misses if CHURN.search(k)],
    }

Only fetch logs for jobs whose steps include a cache step, which the jobs endpoint exposes in each job's steps list. Fetching every log in a 40-repository organization is slow and pointless.

The churn_keys field catches the most common self-inflicted miss: a key that includes the commit SHA or run ID. Every run saves a fresh cache that nothing will ever restore, the repository's cache storage fills, and GitHub evicts the caches that were actually useful. Check the storage side with the CLI:

gh api repos/$GH_ORG/$REPO/actions/cache/usage
gh api repos/$GH_ORG/$REPO/actions/caches --paginate \
  -q '.actions_caches[] | [.size_in_bytes, .last_accessed_at, .ref, .key] | @tsv' \
  | sort -rn | head -20

A repository sitting at its cache size limit with hundreds of entries whose last_accessed_at equals their creation time has a churning key. The agent reports it as a finding, and the fix is always the same shape: key on the lockfile hash and add a prefix fallback.

- uses: actions/cache@v4
  with:
    path: ~/.npm
    key: npm-${{ runner.os }}-${{ hashFiles('**/package-lock.json') }}
    restore-keys: |
      npm-${{ runner.os }}-

Step 4: The Model Writes the Patch, actionlint Checks It

Everything above produces a JSON list of findings per repository with a measured cost. The model receives one repository's findings plus the full text of each workflow file named in them, and returns a patch. It does not get the billing data, the token, or any tool that writes.

You are editing GitHub Actions workflow files to remove measured waste.
You will receive findings (pattern, workflow path, job, measured minutes per 14 days)
and the full text of each workflow file.

Rules:
- Make the minimal change that fixes each finding. Do not refactor, rename, or reorder.
- For SUPERSEDED_RUN add a top-level concurrency block. Never cancel in-progress runs on the
  default branch or on tags; use cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}.
- For NO_TIMEOUT set timeout-minutes to 3x the median duration, rounded up to 5, minimum 10.
- For CACHE_CHURN replace the key with a hashFiles() key over the lockfile and add restore-keys.
- For PLATFORM_MULTIPLIER move the job to ubuntu-latest only if no step uses xcodebuild,
  codesign, or a Windows-only tool. Otherwise return action NEEDS_HUMAN.
- For OVERSIZED_RUNNER propose the next size down and say that the PR must be benchmarked.
- For ROUNDING_TAX do not merge jobs. Return action NEEDS_HUMAN with the merge suggestion.
Return JSON only: a list of objects with fields finding_id, action, workflow_path,
unified_diff, estimated_minutes_saved_per_month, rationale (max 2 sentences).
Allowed actions: ADD_CONCURRENCY, SET_TIMEOUT, FIX_CACHE_KEY, MOVE_TO_LINUX,
DOWNSIZE_RUNNER, NEEDS_HUMAN.

The refusal to merge jobs is on purpose. Restructuring a matrix changes test parallelism and failure attribution, which is a design decision, not a cost fix. The agent surfaces it with the measured tax and leaves the merge to the team.

The patch goes through two deterministic checks before anyone sees it. First, apply it in a scratch checkout and run actionlint, which catches malformed expressions, unknown runner labels, and bad concurrency syntax. Second, reject any diff that touches a line outside the job the finding named, except for the top-level concurrency block.

git apply --check patch.diff && git apply patch.diff
actionlint -color never .github/workflows/*.yml || exit 1
git diff --stat | tail -1   # refuse if more files changed than findings

A patch that passes opens a pull request titled with the saving, carrying the finding table and the exact runs it measured. The same PR-not-push discipline covered in GitOps for AI agents applies here, and the repository's own CI runs on the PR, so a broken workflow fails visibly instead of silently.

Guardrails

Two identities, not one. The reader uses a fine-grained token with read-only access to Actions, metadata, and organization billing. The writer is a GitHub App installation with contents: write, pull_requests: write, and workflows: write. That last permission is the one people miss. The default GITHUB_TOKEN inside a workflow cannot create or modify files under .github/workflows, and a push that tries fails with a refusal naming the missing workflows permission. If you run this agent as a workflow, it needs an App token minted with that permission for the write step only. The GitHub MCP server lockdown post covers the read-only toolset shape if you want the model to pull workflow files through MCP instead of a pre-built packet.

No cache deletion, ever. Deleting caches is a one-line gh cache delete and it is tempting to let the agent do it for churned keys. Do not. A cache miss costs minutes; deleting the one cache a nightly job depends on costs an outage at 03:00. The PR fixes the key, and the old entries age out on their own.

Budgets. One PR per repository per week, no PR for a finding under 100 minutes per month on included-minute plans or under $5 per month on billed plans, and a hard cap on log fetches per run. The model is called once per repository with a bounded packet, so the token cost is predictable and small, the same budgeting that the CI failure triage agent relies on when it runs on every red build.

Measure after merge. Re-run Step 2 on the same repository two weeks after each PR merges and record the delta next to the estimate. Estimates that are consistently wrong in one direction mean a heuristic needs retuning. A team that cannot show the before-and-after will not trust the next PR, and the DORA metrics pipeline you probably already have is the natural place to graph minutes per deploy alongside deployment frequency.

Honest Limits

The oversized-runner heuristic is weak. The API exposes no CPU utilization for hosted runners, so the agent can only observe that a 16-core job's duration barely differs from the same job's history on a smaller label, or that its duration is mostly checkout and dependency install. The PR for this pattern says so, and asks for one benchmark run on the smaller label before merge.

Self-hosted runners show zero in the billing API, so this agent sees nothing on them. The job-level reconstruction in Step 2 still works and still finds superseded runs and missing timeouts, but you have to attach your own price per minute from the node cost.

Log grepping depends on actions/cache keeping its two output lines stable. It has for years, but pin the regex to the version you run and add a test fixture, as the MCP server testing post does for tool output parsing.

And the biggest saving is often one the agent will only describe, never fix: a pipeline that rebuilds a container from scratch on every push because its Dockerfile has no cacheable layers. The agent reports the build job's minutes. Turning them into a multi-stage build is still a human's afternoon.

#ai#devops#github-actions#cicd#finops#cost-optimization#automation
D
DevToCashAuthor

Senior DevOps/SRE Engineer · 10+ years · Professional Trader (IDX, Crypto, US Equities)

I write about real infrastructure patterns and trading strategies I use in production and in live markets. No courses, no affiliate hype — just documentation of what actually works.

More about me →