What This Agent Does
A flaky test quarantine agent collects per-test results from every CI run, finds tests that both passed and failed on the same commit, scores how often each one flips and how many runner minutes it burns, and opens a pull request that moves the worst offenders into a quarantine list. Quarantined tests keep running on every build, but they can no longer fail it. The same data later proves a test has stabilized, and the agent opens the reverse PR. Deterministic code does the detection and the thresholds. The model does one narrow job: read the failure output and the test source, and write a cause hypothesis and a one-line quarantine reason for the human reviewer. The agent never deletes a test, never skips one silently, and never commits to the main branch.
The CI failure triage agent stops exactly where this one starts. That agent refuses to auto-retry a flaky test, because a retried pass buries the flake, and instead names the test so "someone quarantines or fixes it." In most teams that someone never arrives, and the suite slowly turns into a thing people re-run until it is green. This agent is the someone.
What "Flaky" Means in the Data
Only one signal is trustworthy: the same test, on the same commit SHA, with at least one pass and at least one fail. Code did not change, so the outcome depended on something else. Everything weaker (fails on a PR, passes after a rebase) is contaminated by real changes and must not drive a quarantine decision.
| Signature | What produces it | Likely cause |
|---|---|---|
| Pass and fail on the same SHA, same job | Re-run of a failed job, or a nightly stress run | Timing, shared state, external dependency |
| Fails only on one matrix entry | Same SHA, different OS or Python version | Platform or locale assumption |
| Fails only after a specific other test | Visible once you sort runs by execution order | Order dependence, leaked global state |
| Fails only between 23:00 and 01:00 UTC | Cluster of failures at day boundaries | Date arithmetic, time zone assumption |
| Fails on every run of main | Zero passes since a given SHA | Not a flake, a real regression; hand to the triage agent |
That last row is the most important guardrail in the system. A test that fails 100 percent of the time on main is broken, and quarantining it would hide a regression. The scorer refuses to emit a quarantine finding for it.
Step 1: Make CI Produce the Evidence
The agent needs one JUnit XML file per job, uploaded on every run including failures, with any retry plugin switched off in the data path. A plugin like pytest-rerunfailures turns a pass-after-fail into a single green result, which is precisely the flip you need to see.
# .github/workflows/test.yml (excerpt)
- name: Run tests
run: pytest -p no:rerunfailures --junitxml=reports/junit.xml
- name: Upload JUnit report
uses: actions/upload-artifact@v4
if: always()
with:
name: junit-${{ github.job }}-${{ strategy.job-index }}
path: reports/junit.xml
retention-days: 30
Same-SHA data is rare on PR branches, because nobody re-runs a green job. A nightly stress workflow fixes that by running the suite three times on the head of main with random ordering. Three repetitions per night, for a 12-minute suite, is 36 runner minutes a day. Price it against the retries it will remove using the method in the GitHub Actions cost agent.
# .github/workflows/flake-stress.yml
name: flake-stress
on:
schedule:
- cron: "17 3 * * *"
workflow_dispatch:
jobs:
stress:
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
rep: [1, 2, 3]
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- run: pip install -e ".[test]" pytest-randomly
- run: pytest -p randomly -p no:rerunfailures --junitxml=reports/junit.xml
- uses: actions/upload-artifact@v4
if: always()
with:
name: junit-stress-${{ matrix.rep }}
path: reports/junit.xml
Step 2: Collect Per-Test Results Across Runs
The collector walks the last 14 days of workflow runs, downloads every JUnit artifact, and flattens each test case into one row. Quarantined tests will later show up as skipped with the type pytest.xfail, and those rows must count as failures, or the un-quarantine logic will never see the test fail again and will release it too early.
# collect.py — read-only; token needs actions:read on the repo
import io, os, zipfile, requests
import xml.etree.ElementTree as ET
from datetime import datetime, timedelta, timezone
REPO = os.environ["GH_REPO"] # "org/name"
H = {"Authorization": f"Bearer {os.environ['GH_TOKEN']}",
"Accept": "application/vnd.github+json",
"X-GitHub-Api-Version": "2022-11-28"}
API = f"https://api.github.com/repos/{REPO}"
def gh(path, **params):
r = requests.get(f"{API}{path}", headers=H, params=params, timeout=30)
r.raise_for_status()
return r.json()
def runs(days=14):
since = (datetime.now(timezone.utc) - timedelta(days=days)).strftime("%Y-%m-%d")
page = 1
while True:
batch = gh("/actions/runs", created=f">={since}",
per_page=100, page=page)["workflow_runs"]
yield from batch
if len(batch) < 100:
return
page += 1
def junit_rows(run):
for art in gh(f"/actions/runs/{run['id']}/artifacts")["artifacts"]:
if not art["name"].startswith("junit-") or art["expired"]:
continue
z = requests.get(art["archive_download_url"], headers=H, timeout=60)
z.raise_for_status()
with zipfile.ZipFile(io.BytesIO(z.content)) as zf:
for name in zf.namelist():
if not name.endswith(".xml"):
continue
for tc in ET.fromstring(zf.read(name)).iter("testcase"):
sk = tc.find("skipped")
if sk is not None and sk.get("type") != "pytest.xfail":
continue # genuinely skipped, no signal
failed = (sk is not None
or tc.find("failure") is not None
or tc.find("error") is not None)
msg = next((e.get("message", "") for e in tc
if e.tag in ("failure", "error", "skipped")), "")
yield {
"test_id": f"{tc.get('classname')}::{tc.get('name')}",
"matrix": art["name"],
"sha": run["head_sha"], "run_id": run["id"],
"attempt": run["run_attempt"], "branch": run["head_branch"],
"outcome": "fail" if failed else "pass",
"secs": float(tc.get("time") or 0),
"message": msg[:500], "at": run["run_started_at"],
}
rows = [r for run in runs() for r in junit_rows(run)]
The matrix field is part of the test's identity on purpose. A test that fails only on the Windows entry of a matrix is a different finding from one that fails everywhere, and it usually has a different fix.
Step 3: Score Flips, Not Failures
The scorer groups rows by test and matrix entry, counts SHAs on which the test produced both outcomes, and drops anything with no failures or with a 100 percent failure rate on main. No model is involved.
# score.py — deterministic; thresholds are the policy, keep them in code review
import collections, statistics
def score(rows):
by_test = collections.defaultdict(list)
for r in rows:
by_test[(r["test_id"], r["matrix"])].append(r)
out = []
for (test_id, matrix), rs in by_test.items():
fails = [r for r in rs if r["outcome"] == "fail"]
if not fails:
continue
on_main = [r for r in rs if r["branch"] == "main"]
if on_main and all(r["outcome"] == "fail" for r in on_main):
continue # broken, not flaky
by_sha = collections.defaultdict(set)
for r in rs:
by_sha[r["sha"]].add(r["outcome"])
flip_shas = sum(1 for o in by_sha.values() if o == {"pass", "fail"})
out.append({
"test_id": test_id, "matrix": matrix,
"runs": len(rs), "fails": len(fails),
"fail_rate": round(len(fails) / len(rs), 3),
"flip_shas": flip_shas,
"p50_secs": round(statistics.median(r["secs"] for r in rs), 2),
"last_fail": max(fails, key=lambda r: r["at"]),
})
return sorted(out, key=lambda f: (-f["flip_shas"], -f["fail_rate"]))
def should_quarantine(f):
return f["flip_shas"] >= 3 and f["fail_rate"] >= 0.05
def should_release(f):
passes = [r for r in f["rows"] if r["outcome"] == "pass"]
return f["fails"] == 0 and len(passes) >= 25 and len({r["sha"] for r in passes}) >= 10
| Threshold | Value | Why |
|---|---|---|
| Quarantine | 3 or more flip SHAs and fail rate of at least 5 percent in 14 days | One flip is noise; three distinct commits is a pattern |
| Release | 25 consecutive passes across at least 10 distinct SHAs | Proves stability against real code motion, not one lucky night |
| Escalate | Quarantined for more than 30 days | A quarantine list is a to-do list, not a graveyard |
| Weekly budget | At most 5 new quarantines per repository per week | Keeps reviewers reading each PR instead of rubber-stamping |
Every flaky failure on a PR costs more than its own duration. The usual response is a re-run of the whole job, so the wasted minutes are the job's billed minutes multiplied by the number of re-runs the test provoked. Rows with attempt greater than one, where this test was the only failure in attempt one, give you that count directly. Put the number in the PR. A reviewer who sees "this test caused 41 re-runs and 380 billed minutes in two weeks" approves the quarantine in under a minute.
Step 4: Let the Model Write the Hypothesis, Not the Decision
The threshold has already decided. The model's input is the scoring row, the assertion message and the last 40 log lines from the most recent failure, and up to 200 lines of the test source fetched through the contents API. Its output is a structured tool call.
TOOL = {
"name": "propose_quarantine",
"description": "Explain the likely cause of one flaky test and write its quarantine entry.",
"input_schema": {
"type": "object",
"properties": {
"cause": {"type": "string",
"enum": ["timing", "shared_state", "external_dependency",
"resource_pressure", "time_or_locale", "unknown"]},
"confidence": {"type": "number", "minimum": 0, "maximum": 1},
"evidence": {"type": "string",
"description": "Quoted log or source lines that support the cause. Empty if none."},
"suggested_fix": {"type": "string"},
"quarantine_reason": {"type": "string",
"description": "One line, under 80 characters, for quarantine.txt"}
},
"required": ["cause", "confidence", "evidence", "quarantine_reason"]
}
}
The system prompt is short and asymmetric in the same way the triage agent's is:
You explain why a test is flaky. You do not decide whether to quarantine it;
that decision was already made by a threshold and is not yours to revisit.
Choose "unknown" unless the evidence field can quote a specific line that
supports the cause. Never recommend deleting or skipping the test. A wrong
confident hypothesis costs an engineer an afternoon; "unknown" costs nothing.
A sleep(0.5) before an assertion, a module-level dictionary mutated by two tests, a requests.get to a real hostname, or datetime.now() compared against a hard-coded date are the four patterns that account for most hypotheses, and the model is good at spotting them when the source is in front of it. Label the output "hypothesis" in the PR body regardless of the confidence number. The confidence is calibrated against the fixture set from your agent evals, not against reality.
Step 5: Quarantine Mechanically
The quarantine is one text file in the repository and one hook in conftest.py. Listed tests still run and still report, but a failure becomes an expected failure and cannot turn the build red. The identifier format matches the JUnit classname::name pair the collector produces, so the same string flows from data to file to hook without translation.
# tests/quarantine.txt — one test per line: junit_id # reason | opened | owner
tests.api.test_orders::test_refund_idempotent # timing: polls status with no wait | 2026-10-05 | @org/payments
tests.workers.test_scheduler::test_runs_at_midnight # time_or_locale: naive datetime | 2026-10-05 | @org/platform
# tests/conftest.py
import pathlib, pytest
QFILE = pathlib.Path(__file__).parent / "quarantine.txt"
def _quarantined():
out = {}
if QFILE.exists():
for line in QFILE.read_text().splitlines():
line = line.strip()
if line and not line.startswith("#"):
test_id, _, rest = line.partition(" # ")
out[test_id.strip()] = rest.split("|")[0].strip() or "quarantined"
return out
def _junit_id(item):
parts = item.nodeid.split("::")
module = parts[0].removesuffix(".py").replace("/", ".")
return ".".join([module] + parts[1:-1]) + "::" + parts[-1]
def pytest_collection_modifyitems(config, items):
q = _quarantined()
for item in items:
reason = q.get(_junit_id(item))
if reason:
item.add_marker(pytest.mark.xfail(reason=reason, strict=False))
strict=False matters. With strict mode, an unexpected pass would fail the build, which defeats the point. With it off, a quarantined test that passes is simply a pass, and that pass is exactly the data the release rule counts.
The agent writes the line, commits on a branch, and opens the PR. It never has push access to main, for the same reasons laid out in GitOps for AI agents.
# open_pr.sh — called once per quarantine finding that passed the budget check
set -euo pipefail
BR="quarantine/$(date +%Y%m%d)-${TEST_SLUG}"
git switch -c "$BR"
printf '%s # %s | %s | %s\n' "$TEST_ID" "$REASON" "$(date +%F)" "$OWNER" >> tests/quarantine.txt
git commit -am "test: quarantine ${TEST_ID} (flaky, ${FLIP_SHAS} flip SHAs in 14d)"
git push -u origin "$BR"
gh pr create --title "Quarantine flaky test: ${TEST_ID}" \
--body-file "$BODY_FILE" --label flaky-test --reviewer "${OWNER#@}"
The owner comes from CODEOWNERS for the test file's directory, with a fallback to the author of the last commit that touched the test. The PR body carries the evidence table (runs, fails, flip SHAs, re-runs provoked, billed minutes), the last assertion message in a code block, and the model's hypothesis under a heading that says hypothesis.
Step 6: Un-Quarantine, and Escalate What Never Heals
Because quarantined tests keep running, the same collector keeps scoring them. When one meets the release rule, the agent opens a PR that deletes its line. When one has been in the file for more than 30 days without meeting it, the agent opens an issue assigned to the owner with the current flip history and the original hypothesis, titled "fix or delete." A test that has been unable to fail the build for a month is not protecting anything, and the issue says so.
Guardrails That Are Not Optional
| Rule | Enforced where |
|---|---|
| Never quarantine a test with a 100 percent failure rate on main | Scorer, before the model sees it |
| Never quarantine a test modified in the last 3 days | Scorer; someone is already fixing it |
Never quarantine tests under a protected path such as tests/security/ or tests/billing/ | Deny-list in the PR step |
| At most 5 new quarantine PRs per repository per week | Budget counter keyed on the PR label |
| PR only; the agent's token cannot push to main | Branch protection plus token scope |
| Model output is labelled hypothesis and never changes the decision | Prompt and PR template |
Honest Limits
Retry plugins anywhere in the pipeline poison the data. If a team cannot turn them off in the main workflow, run the collector only on the stress workflow, where you control the flags.
Parametrized tests share a module and class but differ in the bracketed parameter in the name, so they are scored separately. That is correct for most flakes and wrong for a shared fixture that flakes for every parameter. Group by test name without the bracket when the flip set covers more than half the parameters.
The flake rate on PR branches is inflated by real breakage in progress. Weight the stress runs and main-branch runs higher than PR runs when thresholds are close, or simply compute the fail rate from main and stress rows only and use PR rows just for the re-run cost figure.
Finally, this agent changes your DORA numbers. A flaky suite inflates change failure rate with failures that were never changes. Once quarantines land, the DORA metrics pipeline shows change failure rate dropping, and the deployment risk scorer stops penalizing services for noise. Measure the suite-wide flake rate (flipping tests divided by total tests, weekly) and put it next to those numbers, because the moment it stops falling is the moment the quarantine file has become a graveyard again.