AI agents moved our bottleneck from writing code to trusting it and landing it. How we rebuilt our delivery pipeline for an agent fleet, with the numbers.

Steve Saper
Founder & CEO of DeepWind. 15 years advising Fortune 500 product teams before building the product operating system for the AI era -- closing the gap between customer signal and what actually gets shipped. Also hosts The Product Briefing.

AI agents let our small team ship at a pace no traditional team could: one engineer directing agents shipped 1,000 story points in a single sprint. That speed moved our bottleneck from writing code to trusting it and landing it. This is how our delivery pipeline had to grow up, what we chose and why, and what it means for teams adopting agentic workflows with DeepWind.
We built DeepWind agent-first from day one. A handful of engineers direct a fleet of AI coding agents: builders take a brief and write the change, separate reviewer agents check it, and coordinator agents sequence the work.
Early on, the pipeline behind that was simple. An agent opened a PR, tests ran, the PR merged, and every merge auto-deployed to staging, with frequent production releases behind it. With a small app and a few agents, that loop was fast and good enough.
Then three things grew at once: the product (hundreds of routes, a multi-tenant database with row-level security, real-time agent features), the test suite that protected it, and the number of agents opening PRs at the same moment.

The bars aren't like for like, and that is the point. Before May 23, agents committed straight to main, so each bar counts raw commits. After it, every change is a reviewed, squash-merged PR carrying far more work per landing.
By September, writing code was no longer the constraint. Getting it safely onto main was.
The failures were rarely flaky tests. On one measured day, all 20 failed runs were deterministic. The cost was work that was never going to count: runs on stale code, runs judged by an outdated copy of the checks, runs that raced a newer merge.
We chose external sandboxes for scale, a dedicated machine for the few changes that alter the gate itself, and an agent-specific trust layer on top of both.
| Option | What it offered | Why it wasn't enough on its own |
|---|---|---|
| Buy more hosted CI minutes | Least effort | Cost grows with every agent; no protection against agents weakening their own checks |
| Self-hosted runners on our own machines | Full control, no per-minute cost | Capacity capped by hardware; one busy machine slows every PR |
| An off-the-shelf merge queue | Orders merges automatically | Solves ordering, not trust or collisions between autonomous agents |
| Slow the agents down | Fewer collisions | Throws away the velocity that is the whole point |
| External sandboxes plus our own trust layer (chosen) | Parallel scale-out, isolation, per-run evidence | Needed real engineering: receipts, pinned checks, coordination |
In hindsight, we built some plumbing that mature products already offer. The part worth building was the layer that makes agent-written code trustworthy at speed.
External sandboxes turned gating from a queue into a fan-out: every PR gets its own disposable machine with a fresh database, and a dozen can run at once for about $0.23 each.

main, so a PR can't bring its own lenient copy.Eight pieces carry most of the weight. Each exists because a specific failure cost us real time.
| Tool | What it does | The problem it solves |
|---|---|---|
| Sandbox gate runner | Runs the full quality gate for one PR in a disposable cloud sandbox | Parallel capacity without owning hardware |
| Gate receipts | Records a green result bound to the exact commit, base and version of the checks | Stale or mismatched results can't authorize a merge |
| Pinned checks | Every pass/fail script runs from main, never from the branch under test | An agent can't weaken its own checks, even by accident |
| Independent review attestation | A separate agent reviews the exact commit; the verdict is a merge prerequisite | The author never approves its own work |
| Guarded merge | The only path to main: checks the receipt, the review and whether main moved | No bypasses, and a clear reason whenever it refuses |
| Speculative merge queue | Gates each PR against the latest main and lands them in order | Fewer re-runs when main moves under a PR |
| Step reuse and affected-step re-runs | Reuses results for steps a change can't affect; re-runs only the rest | A small change stops paying for a full re-run |
| Preflight doctor | Predicts in seconds whether a run can ever be used | Stops hour-long runs that were doomed at the start |
Around them sit a ratchet that lets known failures stay known without letting new ones in, and a disposable real database per run so security tests exercise real row-level rules.
The lesson was not "test more". It was: make every check trustworthy, and never spend an hour on a run that could never count.
The results so far: median time-to-merge fell from 8.1 hours to 2.3 hours, and gate runs per merged PR from 6.5 to 2.9. The preflight doctor and affected-step re-runs are landing now, and we measure a full week from September 28.

Flaky tests were never the main problem. Most failed runs were predictable: stale code, an outdated copy of the checks, or real defects that a quick preflight or a local test run would have caught.
Gating more work took more runs. In four weeks, merged PRs went from about 2 sandbox gate runs each to about 5. Most of that growth is new revisions of the same PR, not reruns of the same code: the number of distinct commits gated per merged PR rose from 1.7 to 3.4, as agents refreshed branches onto a moving main and fixed what the gate found.

The waiting around the gate shrank even as gated merges grew ninefold, from 14 a week to 131. A new PR now reaches its first gate run in a median of 34 minutes, down from about 9 hours, and nine in ten green PRs merge within about an hour of going green, down from 12 hours.

The cost moved into the gate itself. A single run now takes a median of 28 minutes, up from 8, because the test suite and the list of checks grew with the product. Cutting that is our next target: rerun only the checks a change can affect, instead of everything.

The week of September 21 felt faster, and the numbers agree. Compared with the week before:
| Measure | Week of Sep 14 | Week of Sep 21 | Change |
|---|---|---|---|
| PRs merged | 134 | 142 | +6% |
| Merged within a day of opening | 82 | 107 | +30% |
| Median time from PR opened to merge (PRs merged that week) | 15.3 h | 5.3 h | about 3x faster |
| Failed gate runs | 64.9% | 45.8% | 19 points lower |
| Median time from PR opened to first gate run | 64 min | 34 min | about half |
| Green gate to merge, 90th percentile | 7.6 h | 1.0 h | about 7x faster |
| Gate runs per merged PR (average) | 4.7 | 5.2 | worse |
| Median length of one gate run | 25 min | 28 min | worse |
Most of the gain came from removing reasons to fail, not from adding checks:
Two numbers got worse: each merged PR still takes about five gate runs, and each run is a little longer. That is where the next round of work goes. September 28 is not in the comparison: we paused merges for most of that day to land a batch of gate changes, so its numbers reflect the pause, not the pipeline.
Mostly yes. The two numbers that went the wrong way measure effort, not waiting, and most of the extra effort is the gate doing its job.
Why each PR takes more runs. A gate result counts only for the exact code it checked, on top of the exact main it will land on. With about 20 merges a day, main moves every hour or so. Any PR open for a few hours has to catch up with main before it lands, and that new version is checked again. Agents also push fixes for what the gate finds. Both create a new version of the PR, and every version gets gated. In the week of September 21:
Why each run is longer. Most of the growth happened between late August and mid-September, when a run went from 8 to 25 minutes. That is when the sandbox took over checks that used to run only on a developer machine: the tests that need a real database, and the gate's own self-tests. Those local runs could take close to an hour and tied up the machine while they ran. Since then the length has held roughly steady, 25 to 28 minutes, while the product and its test suite kept growing.
Why that trade is acceptable. Sandbox runs happen in parallel and nobody sits waiting for them: agents move on to other work while the gate runs. Compute is cheap next to engineering time. What matters for delivery is how long a finished change waits to land, and that fell sharply (see the table above).
Where it is not fine. Two costs are real:
September 28 gave a clear example. A billing change we needed for a launch failed the gate three times, and none of the failures came from the change itself:
Chasing these also turned up a test that was quietly broken on main and would have failed some other unlucky PR next.
What we are doing about it:
The goal is fewer runs per change, not fewer checks.
We priced every model call our agents made, from their session logs at list prices, and every sandbox gate run, and divided by the PRs merged each week:
| Week of | PRs merged | Agent model cost | Sandbox gate cost | Cost per merged PR |
|---|---|---|---|---|
| Aug 17 | 109 | $2,762 | not yet in use | $25 |
| Aug 24 | 113 | $8,164 | not yet in use | $72 |
| Aug 31 | 164 | $10,248 | $3 | $62 |
| Sep 7 | 73 | $7,363 | $19 | $101 |
| Sep 14 | 134 | $10,723 | $104 | $81 |
| Sep 21 | 142 | $10,684 | $128 | $76 |
This is a fully loaded figure: it includes every agent session that week, including planning, reviews, marketing and the work of running the pipeline itself, not just the work of writing each change. Before mid-August our session logs are incomplete, so we leave those weeks out.
Two things stand out. First, the sandbox gate is about 1% of the bill: an extra gate run costs cents. The expensive part of a false failure is not the rerun but the agent time spent investigating it, which is why removing false failures matters more than making runs cheaper. Second, cost per PR has come down from its September 7 peak while volume doubled, which is the same trend as the wait times above.
If your team is moving to agentic workflows, you will hit the same wall, usually sooner than you expect. More agents means more PRs at once, more collisions, and more chances for an agent to pass a check without meeting its intent.
We run DeepWind the way we ask you to run your delivery, so what we learned goes straight into the product:
The practical effect for customers: fixes and features reach you faster, and they have been through the same exact-commit review and pinned checks described above before they ship.
You can start without paying: the free tools give an agent-driven team structure on day one, and a subscription closes the loop across your whole portfolio.
| Free: DeepWind Harness | Subscribed: DeepWind | |
|---|---|---|
| What it is | A public, versioned workflow bundle for Claude Code and Codex, from the open DeepWindAI/harness repo | The DeepWind app, connected to your agents through the DeepWind MCP |
| What you get | Agent roles, skills and frameworks; a merge gate that stops an agent from self-merging an unreviewed, security-sensitive PR; the gating practices in this post | Briefs tied to objectives, delivery forecasts, agent fleet metrics, and drift caught and re-aimed every cycle |
| What it changes | Your agents follow a disciplined build, review and merge flow from day one | Each cycle is aimed at what actually moved your number |
| How to start | curl -fsSL https://deepwind.ai/install | bash | Start at deepwind.ai/start with $30 in free credit, no card needed |
The harness is useful on its own because the hard part of agentic delivery is structure, not tooling: clear briefs, independent review, and checks an agent can't bend. We are adding this gating work to it, including the preflight check that stops doomed runs, so every team gets it for free. A subscription closes the loop: it shows which of that work is moving your objectives, and aims the next cycle there.