devops

Build a Flaky Test Quarantine Agent: Score Flake Rates from JUnit History, Quarantine with a PR, and Un-Quarantine When Tests Stabilize

Build a flaky test quarantine agent: score flake rates from JUnit history across CI runs, quarantine with a reviewed PR, and un-quarantine once tests stabilize.

October 5, 2026·13 min read·
#ai#devops#cicd#github-actions#automation#platform-engineering

What This Agent Does

A flaky test quarantine agent collects per-test results from every CI run, finds tests that both passed and failed on the same commit, scores how often each one flips and how many runner minutes it burns, and opens a pull request that moves the worst offenders into a quarantine list. Quarantined tests keep running on every build, but they can no longer fail it. The same data later proves a test has stabilized, and the agent opens the reverse PR. Deterministic code does the detection and the thresholds. The model does one narrow job: read the failure output and the test source, and write a cause hypothesis and a one-line quarantine reason for the human reviewer. The agent never deletes a test, never skips one silently, and never commits to the main branch.

The CI failure triage agent stops exactly where this one starts. That agent refuses to auto-retry a flaky test, because a retried pass buries the flake, and instead names the test so "someone quarantines or fixes it." In most teams that someone never arrives, and the suite slowly turns into a thing people re-run until it is green. This agent is the someone.

What "Flaky" Means in the Data

Only one signal is trustworthy: the same test, on the same commit SHA, with at least one pass and at least one fail. Code did not change, so the outcome depended on something else. Everything weaker (fails on a PR, passes after a rebase) is contaminated by real changes and must not drive a quarantine decision.

SignatureWhat produces itLikely cause
Pass and fail on the same SHA, same jobRe-run of a failed job, or a nightly stress runTiming, shared state, external dependency
Fails only on one matrix entrySame SHA, different OS or Python versionPlatform or locale assumption
Fails only after a specific other testVisible once you sort runs by execution orderOrder dependence, leaked global state
Fails only between 23:00 and 01:00 UTCCluster of failures at day boundariesDate arithmetic, time zone assumption
Fails on every run of mainZero passes since a given SHANot a flake, a real regression; hand to the triage agent

That last row is the most important guardrail in the system. A test that fails 100 percent of the time on main is broken, and quarantining it would hide a regression. The scorer refuses to emit a quarantine finding for it.

Step 1: Make CI Produce the Evidence

The agent needs one JUnit XML file per job, uploaded on every run including failures, with any retry plugin switched off in the data path. A plugin like pytest-rerunfailures turns a pass-after-fail into a single green result, which is precisely the flip you need to see.

# .github/workflows/test.yml (excerpt)
- name: Run tests
  run: pytest -p no:rerunfailures --junitxml=reports/junit.xml
- name: Upload JUnit report
  uses: actions/upload-artifact@v4
  if: always()
  with:
    name: junit-${{ github.job }}-${{ strategy.job-index }}
    path: reports/junit.xml
    retention-days: 30

Same-SHA data is rare on PR branches, because nobody re-runs a green job. A nightly stress workflow fixes that by running the suite three times on the head of main with random ordering. Three repetitions per night, for a 12-minute suite, is 36 runner minutes a day. Price it against the retries it will remove using the method in the GitHub Actions cost agent.

# .github/workflows/flake-stress.yml
name: flake-stress
on:
  schedule:
    - cron: "17 3 * * *"
  workflow_dispatch:
jobs:
  stress:
    runs-on: ubuntu-latest
    strategy:
      fail-fast: false
      matrix:
        rep: [1, 2, 3]
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: "3.12"
      - run: pip install -e ".[test]" pytest-randomly
      - run: pytest -p randomly -p no:rerunfailures --junitxml=reports/junit.xml
      - uses: actions/upload-artifact@v4
        if: always()
        with:
          name: junit-stress-${{ matrix.rep }}
          path: reports/junit.xml

Step 2: Collect Per-Test Results Across Runs

The collector walks the last 14 days of workflow runs, downloads every JUnit artifact, and flattens each test case into one row. Quarantined tests will later show up as skipped with the type pytest.xfail, and those rows must count as failures, or the un-quarantine logic will never see the test fail again and will release it too early.

# collect.py — read-only; token needs actions:read on the repo
import io, os, zipfile, requests
import xml.etree.ElementTree as ET
from datetime import datetime, timedelta, timezone

REPO = os.environ["GH_REPO"]            # "org/name"
H = {"Authorization": f"Bearer {os.environ['GH_TOKEN']}",
     "Accept": "application/vnd.github+json",
     "X-GitHub-Api-Version": "2022-11-28"}
API = f"https://api.github.com/repos/{REPO}"

def gh(path, **params):
    r = requests.get(f"{API}{path}", headers=H, params=params, timeout=30)
    r.raise_for_status()
    return r.json()

def runs(days=14):
    since = (datetime.now(timezone.utc) - timedelta(days=days)).strftime("%Y-%m-%d")
    page = 1
    while True:
        batch = gh("/actions/runs", created=f">={since}",
                   per_page=100, page=page)["workflow_runs"]
        yield from batch
        if len(batch) < 100:
            return
        page += 1

def junit_rows(run):
    for art in gh(f"/actions/runs/{run['id']}/artifacts")["artifacts"]:
        if not art["name"].startswith("junit-") or art["expired"]:
            continue
        z = requests.get(art["archive_download_url"], headers=H, timeout=60)
        z.raise_for_status()
        with zipfile.ZipFile(io.BytesIO(z.content)) as zf:
            for name in zf.namelist():
                if not name.endswith(".xml"):
                    continue
                for tc in ET.fromstring(zf.read(name)).iter("testcase"):
                    sk = tc.find("skipped")
                    if sk is not None and sk.get("type") != "pytest.xfail":
                        continue                       # genuinely skipped, no signal
                    failed = (sk is not None
                              or tc.find("failure") is not None
                              or tc.find("error") is not None)
                    msg = next((e.get("message", "") for e in tc
                                if e.tag in ("failure", "error", "skipped")), "")
                    yield {
                        "test_id": f"{tc.get('classname')}::{tc.get('name')}",
                        "matrix": art["name"],
                        "sha": run["head_sha"], "run_id": run["id"],
                        "attempt": run["run_attempt"], "branch": run["head_branch"],
                        "outcome": "fail" if failed else "pass",
                        "secs": float(tc.get("time") or 0),
                        "message": msg[:500], "at": run["run_started_at"],
                    }

rows = [r for run in runs() for r in junit_rows(run)]

The matrix field is part of the test's identity on purpose. A test that fails only on the Windows entry of a matrix is a different finding from one that fails everywhere, and it usually has a different fix.

Step 3: Score Flips, Not Failures

The scorer groups rows by test and matrix entry, counts SHAs on which the test produced both outcomes, and drops anything with no failures or with a 100 percent failure rate on main. No model is involved.

# score.py — deterministic; thresholds are the policy, keep them in code review
import collections, statistics

def score(rows):
    by_test = collections.defaultdict(list)
    for r in rows:
        by_test[(r["test_id"], r["matrix"])].append(r)

    out = []
    for (test_id, matrix), rs in by_test.items():
        fails = [r for r in rs if r["outcome"] == "fail"]
        if not fails:
            continue
        on_main = [r for r in rs if r["branch"] == "main"]
        if on_main and all(r["outcome"] == "fail" for r in on_main):
            continue                                   # broken, not flaky
        by_sha = collections.defaultdict(set)
        for r in rs:
            by_sha[r["sha"]].add(r["outcome"])
        flip_shas = sum(1 for o in by_sha.values() if o == {"pass", "fail"})
        out.append({
            "test_id": test_id, "matrix": matrix,
            "runs": len(rs), "fails": len(fails),
            "fail_rate": round(len(fails) / len(rs), 3),
            "flip_shas": flip_shas,
            "p50_secs": round(statistics.median(r["secs"] for r in rs), 2),
            "last_fail": max(fails, key=lambda r: r["at"]),
        })
    return sorted(out, key=lambda f: (-f["flip_shas"], -f["fail_rate"]))

def should_quarantine(f):
    return f["flip_shas"] >= 3 and f["fail_rate"] >= 0.05

def should_release(f):
    passes = [r for r in f["rows"] if r["outcome"] == "pass"]
    return f["fails"] == 0 and len(passes) >= 25 and len({r["sha"] for r in passes}) >= 10
ThresholdValueWhy
Quarantine3 or more flip SHAs and fail rate of at least 5 percent in 14 daysOne flip is noise; three distinct commits is a pattern
Release25 consecutive passes across at least 10 distinct SHAsProves stability against real code motion, not one lucky night
EscalateQuarantined for more than 30 daysA quarantine list is a to-do list, not a graveyard
Weekly budgetAt most 5 new quarantines per repository per weekKeeps reviewers reading each PR instead of rubber-stamping

Every flaky failure on a PR costs more than its own duration. The usual response is a re-run of the whole job, so the wasted minutes are the job's billed minutes multiplied by the number of re-runs the test provoked. Rows with attempt greater than one, where this test was the only failure in attempt one, give you that count directly. Put the number in the PR. A reviewer who sees "this test caused 41 re-runs and 380 billed minutes in two weeks" approves the quarantine in under a minute.

Step 4: Let the Model Write the Hypothesis, Not the Decision

The threshold has already decided. The model's input is the scoring row, the assertion message and the last 40 log lines from the most recent failure, and up to 200 lines of the test source fetched through the contents API. Its output is a structured tool call.

TOOL = {
    "name": "propose_quarantine",
    "description": "Explain the likely cause of one flaky test and write its quarantine entry.",
    "input_schema": {
        "type": "object",
        "properties": {
            "cause": {"type": "string",
                      "enum": ["timing", "shared_state", "external_dependency",
                               "resource_pressure", "time_or_locale", "unknown"]},
            "confidence": {"type": "number", "minimum": 0, "maximum": 1},
            "evidence": {"type": "string",
                         "description": "Quoted log or source lines that support the cause. Empty if none."},
            "suggested_fix": {"type": "string"},
            "quarantine_reason": {"type": "string",
                                  "description": "One line, under 80 characters, for quarantine.txt"}
        },
        "required": ["cause", "confidence", "evidence", "quarantine_reason"]
    }
}

The system prompt is short and asymmetric in the same way the triage agent's is:

You explain why a test is flaky. You do not decide whether to quarantine it;
that decision was already made by a threshold and is not yours to revisit.
Choose "unknown" unless the evidence field can quote a specific line that
supports the cause. Never recommend deleting or skipping the test. A wrong
confident hypothesis costs an engineer an afternoon; "unknown" costs nothing.

A sleep(0.5) before an assertion, a module-level dictionary mutated by two tests, a requests.get to a real hostname, or datetime.now() compared against a hard-coded date are the four patterns that account for most hypotheses, and the model is good at spotting them when the source is in front of it. Label the output "hypothesis" in the PR body regardless of the confidence number. The confidence is calibrated against the fixture set from your agent evals, not against reality.

Step 5: Quarantine Mechanically

The quarantine is one text file in the repository and one hook in conftest.py. Listed tests still run and still report, but a failure becomes an expected failure and cannot turn the build red. The identifier format matches the JUnit classname::name pair the collector produces, so the same string flows from data to file to hook without translation.

# tests/quarantine.txt — one test per line: junit_id  # reason | opened | owner
tests.api.test_orders::test_refund_idempotent  # timing: polls status with no wait | 2026-10-05 | @org/payments
tests.workers.test_scheduler::test_runs_at_midnight  # time_or_locale: naive datetime | 2026-10-05 | @org/platform
# tests/conftest.py
import pathlib, pytest

QFILE = pathlib.Path(__file__).parent / "quarantine.txt"

def _quarantined():
    out = {}
    if QFILE.exists():
        for line in QFILE.read_text().splitlines():
            line = line.strip()
            if line and not line.startswith("#"):
                test_id, _, rest = line.partition("  # ")
                out[test_id.strip()] = rest.split("|")[0].strip() or "quarantined"
    return out

def _junit_id(item):
    parts = item.nodeid.split("::")
    module = parts[0].removesuffix(".py").replace("/", ".")
    return ".".join([module] + parts[1:-1]) + "::" + parts[-1]

def pytest_collection_modifyitems(config, items):
    q = _quarantined()
    for item in items:
        reason = q.get(_junit_id(item))
        if reason:
            item.add_marker(pytest.mark.xfail(reason=reason, strict=False))

strict=False matters. With strict mode, an unexpected pass would fail the build, which defeats the point. With it off, a quarantined test that passes is simply a pass, and that pass is exactly the data the release rule counts.

The agent writes the line, commits on a branch, and opens the PR. It never has push access to main, for the same reasons laid out in GitOps for AI agents.

# open_pr.sh — called once per quarantine finding that passed the budget check
set -euo pipefail
BR="quarantine/$(date +%Y%m%d)-${TEST_SLUG}"
git switch -c "$BR"
printf '%s  # %s | %s | %s\n' "$TEST_ID" "$REASON" "$(date +%F)" "$OWNER" >> tests/quarantine.txt
git commit -am "test: quarantine ${TEST_ID} (flaky, ${FLIP_SHAS} flip SHAs in 14d)"
git push -u origin "$BR"
gh pr create --title "Quarantine flaky test: ${TEST_ID}" \
  --body-file "$BODY_FILE" --label flaky-test --reviewer "${OWNER#@}"

The owner comes from CODEOWNERS for the test file's directory, with a fallback to the author of the last commit that touched the test. The PR body carries the evidence table (runs, fails, flip SHAs, re-runs provoked, billed minutes), the last assertion message in a code block, and the model's hypothesis under a heading that says hypothesis.

Step 6: Un-Quarantine, and Escalate What Never Heals

Because quarantined tests keep running, the same collector keeps scoring them. When one meets the release rule, the agent opens a PR that deletes its line. When one has been in the file for more than 30 days without meeting it, the agent opens an issue assigned to the owner with the current flip history and the original hypothesis, titled "fix or delete." A test that has been unable to fail the build for a month is not protecting anything, and the issue says so.

Guardrails That Are Not Optional

RuleEnforced where
Never quarantine a test with a 100 percent failure rate on mainScorer, before the model sees it
Never quarantine a test modified in the last 3 daysScorer; someone is already fixing it
Never quarantine tests under a protected path such as tests/security/ or tests/billing/Deny-list in the PR step
At most 5 new quarantine PRs per repository per weekBudget counter keyed on the PR label
PR only; the agent's token cannot push to mainBranch protection plus token scope
Model output is labelled hypothesis and never changes the decisionPrompt and PR template

Honest Limits

Retry plugins anywhere in the pipeline poison the data. If a team cannot turn them off in the main workflow, run the collector only on the stress workflow, where you control the flags.

Parametrized tests share a module and class but differ in the bracketed parameter in the name, so they are scored separately. That is correct for most flakes and wrong for a shared fixture that flakes for every parameter. Group by test name without the bracket when the flip set covers more than half the parameters.

The flake rate on PR branches is inflated by real breakage in progress. Weight the stress runs and main-branch runs higher than PR runs when thresholds are close, or simply compute the fail rate from main and stress rows only and use PR rows just for the re-run cost figure.

Finally, this agent changes your DORA numbers. A flaky suite inflates change failure rate with failures that were never changes. Once quarantines land, the DORA metrics pipeline shows change failure rate dropping, and the deployment risk scorer stops penalizing services for noise. Measure the suite-wide flake rate (flipping tests divided by total tests, weekly) and put it next to those numbers, because the moment it stops falling is the moment the quarantine file has become a graveyard again.

#ai#devops#cicd#github-actions#automation#platform-engineering
D
DevToCashAuthor

Senior DevOps/SRE Engineer · 10+ years · Professional Trader (IDX, Crypto, US Equities)

I write about real infrastructure patterns and trading strategies I use in production and in live markets. No courses, no affiliate hype — just documentation of what actually works.

More about me →