>the work

These are the full write-ups behind the numbers on the home page, including an interactive explorer of the real 45-ticket drain graph and the pipeline from intake to production.

##case-studies

Production work, written up in terms that travel outside the company that paid for it.

[01] loops-and-graphs

##prototype-to-drained-epic

The highest-value and most complex feature I have shipped so far, followed all the way through: the pipeline ingested a product designer's prototype, decomposed it into a versioned spec, and planned a Linear epic of 45 dependency-wired tickets, which parallel agent lanes drained in about 35 operator-paced hours. I presented the whole thing publicly as "Agentic Loops and Graphs" at Planet DDS AI Meetup #3.

44/45
slices shipped & merged
~35h
first → last merge
~92
PRs across 3 repos
114
requirements authored first
111/111
e2e scenarios green
0
reverts / broken main

The funnel started with parallel readers mapping the prototype against the shipped specs, and they surfaced something I had not seen, which is that the prototype quietly bundled three separate initiatives. The highest-leverage human decision of the entire project was the scope gate right there, where I deleted roughly two epics of accidental scope before any code existed. From there we reconciled 122 requirement candidates against the shipped catalogs, which turned up 19 conflicts that were all resolved and 6 product questions I had to rule on myself, and then we authored 114 requirements with permanent, immutable IDs, all before a single line of implementation.

Then the spec was minted into 45 thin, vertical, end-to-end slices wired together by blocked-by edges. Once that graph existed, readiness was something the loop could compute, so it derived the ready set directly from Linear and nobody hand-ordered the work. Laid out by build depth the columns of the graph are waves, and every slice in a column became buildable the moment the wave before it merged. The tall columns are exactly where two lanes ran in parallel, and the critical path set the floor on how fast any of this could go.

Every lane ran the same nine-step spine: pre-flight, sync real data, implement with a file-domain-isolated strike team, verify against the running app, multi-lens review, fix every finding, iterate until clean, capture learnings, and honor the autonomy dial. I tiered verification to risk rather than to whatever was convenient, so the drain lead re-ran every gate personally, read every diff, drove the live app through browser automation, and queried the database directly for ground truth on the write paths. We finished with 111 of 111 end-to-end scenarios green. The defects had already been caught upstream in each lane's own QA loop, so by the time we reached the final gate it was really acting as confirmation.

The human budget for all 45 tickets came to eight decisions: a scope gate, a reversal blessing, six product rulings, the ticket-mint authorization, the drain gate, pacing, and a set of mid-build product questions that I recorded for a PM rather than letting an agent decide them. Every one of the roughly 92 PRs was approved and merged by a human after the artifact-level gates. I am happy to give agents autonomy over execution, and I keep the judgment calls with people.

// the drain, replayed — sanitized interactive record

This is the real 45-ticket dependency graph from that Linear epic, laid out by build depth, so every column is a wave and each slice in it became buildable the moment the wave before it merged. The tall columns are exactly where the loop ran lanes in parallel, and the critical path of 22 hops is the floor. Click a slice to trace what it waited on and what it unblocked. I have scrubbed the identifiers and titles, but the topology, the risk tiers, the verification depth, and the merge timeline are all real.

WS-A · foundation flagWS-B · core write pathWS-C · limits & overridesWS-D · lifecycle & schedulingWS-E · single-item flowWS-F · bulk changeWS-G · bulk addWS-H · enablementH heavy riskL lightX human-owned
wave 0
wave 1
wave 2
wave 3
wave 4
wave 5
wave 6
wave 7
wave 8
wave 9
wave 10
wave 11
wave 12
wave 13
wave 14
wave 15
wave 16
wave 17
wave 18
wave 19
wave 20
wave 21
wave 22
every fullstack merge · jul 13 15:17 → jul 15 01:59
Jul 14Jul 15operator pauseoperator pauseABCDEFGHS-01 — backend feature flag + mode resolution + API gate (07-13 15:17)S-02 — frontend gate + zero-change guard (07-13 15:53)S-03 — drawer → full-page host + back-nav (07-13 17:01)S-04 — durable intent store + API contract (07-13 17:01)S-05 — self-service change modal (07-13 18:02)S-06 — immediate apply path — re-base + resume (07-13 19:03)S-07 — cycle-boundary deferred executor (07-13 20:37)S-08 — one-per-cycle / supersede rules + display (07-13 21:49)S-10 — durable override record + API (07-13 22:54)S-11 — five-state status row (07-14 03:40)S-12 — edit modals + validation (07-14 04:35)S-13 — effective-threshold stop predicate (07-14 05:26)S-14 — reset to org default (07-14 05:06)S-15 — manual-state record + precedence resolver + audit (07-14 03:40)S-16 — shared status-chip taxonomy (07-14 04:35)S-17 — state modal — two modes (07-14 05:53)S-18 — reverse modal + success banner (07-14 11:13)S-19 — auto-resume scheduler (07-14 14:59)S-20 — edit-date modal + write path (07-14 11:13)S-21 — page controls + live status summary (07-14 14:59)S-22 — per-item manage tab (07-14 16:13)S-23 — two sub-lists + add-list derivation (07-14 16:13)S-24 — add modal + adaptive copy (07-14 17:17)S-25 — confirm → additive enable write (07-14 18:41)S-26 — post-add optimistic reflection (07-14 19:20)S-27 — checkbox + FAB scaffolding (shared bulk infra) (07-14 20:03)S-28 — shared 2-step shell + stepper (07-14 20:32)S-29 — step-1 selection chips + cards (07-14 20:58)S-30 — bulk classification engine (07-14 20:49)S-31 — grouped review + per-item exclusions (07-14 21:21)S-32 — bulk apply write path (07-14 22:12)S-33 — bulk success screen + totals (07-14 22:44)S-34 — bulk-add entry + gates (shared-infra fork) (07-14 21:29)S-35 — step-1 availability derivation (07-14 22:00)S-36 — catalog allowlist + eligibility gates (07-14 22:58)S-37 — bulk-add review breakdown (07-14 23:21)S-38 — bulk-add on-confirm writes (07-15 00:00)S-39 — bulk-add success + outcome report (07-15 00:17)S-40 — cohort tab — active / available (07-14 23:28)S-41 — shared enable service (audit-and-pin) (07-15 00:35)S-42 — prerequisite endpoints + pick modal (07-15 01:26)S-43 — pool attach + metering (audit-and-pin) (07-15 01:11)S-44 — bulk enable from item tab (07-15 01:59)S-45 — edge semantics (audit-and-pin) (07-15 01:50)first merge — S-01last merge — S-44 · epic drained

44 of the 45 slices shipped and merged. The one exclusion, S-09, was ruled human-owned at the drain gate and was deliberately wired to gate nothing. The gaps in the timeline are me pacing the work, and because the loop is reconcile-first, resuming after a gap costs nothing.

// outcome

What I came away with is a repeatable pipeline rather than one lucky run: prototype to spec to ticket graph to drain, with human judgment concentrated into eight named gates and every merge backed by verification somebody actually observed. I spent far less of my time on the loop itself than on the graph, the gates, and the memory around it, and that surrounding machinery is what I would carry to the next team.

[02] suite

##compound-engineering-suite

The engine underneath the drain: 15 compound commands, 6 spec commands, and a deterministic debate engine, all sharing one build spine, and since September 2026 running from intake all the way through the production deploy and the soak watch after it.

37%
QA re-open rate before test-cases
~15%
QA re-open rate after
7/7
reopened tickets caught in blind re-derivation
17
requirement-analysis areas per ticket
5
durable artifacts per run
0
pipeline approvals an agent can grant

Two pipelines share one engine, and the triage front door in front of them accepts the work in whatever form it arrives — a prototype, a Figma file, a ticket, a meeting transcript, a freeform prompt. Small work runs a single lane to a PR, and big features go through the spec-decomposition funnel and end in an epic drain. Whether a lane is running solo or as part of a drain, it runs the same spine: real-data sync, a test-case ledger gate, a file-domain-isolated strike team with a QA loop, live verification against the running app, multi-lens read-only review, fix-every-finding iteration, and learnings capture before ship.

The newest stage sits between triage and planning, and it came out of a retrospective. A 75-ticket epic had drained clean, every lane green, and then QA reopened 37% of its tickets after merge. When I traced the findings, 22 of the 43 attributable ones sat in coverage categories that QA's own derivation rules name and that my pre-PR verification had never named, and about half of everything QA found was a spec gap or a ruling rather than a code defect. The tickets had carried no test cases, so verify had been executing its own guesses. So I split QA's ticket-testing skill in two and pulled the derivation half left of the build. It reads the ticket with every comment and attachment, runs a 17-area requirement analysis, and derives cases under a shared rules file whose version is pinned into every ledger, so the pre-build and post-merge derivations cannot drift apart. Rules generate scenarios and only spec rows may supply expected outcomes, so a rule with no governing row becomes a typed spec-gap line with a proposed value, two rows that disagree become a conflict line that is never buildable as-is, and every one of those lines needs a human disposition before an implementer may start. Each case gets a content-hashed ID so a refresh keeps the unchanged ones, a substrate tag that says whether a unit run, an API call, a real browser, or a human can prove it, and a dependency list so a superseded spec row invalidates the cases that leaned on it. The ledger is then executed by the implementer's QA teammate, by verify in a real browser, and finally by QA after merge, all from the same list, which is posted once as a single comment on the ticket. As the acceptance test I re-ran the derivation blind on seven reopened tickets with their inputs frozen to pre-build state, and it reproduced the later QA finding on all seven. The re-open rate since has been about 15%.

Before anything executes, the plan has to survive an argument. The debate engine runs a 5-lens reviewer panel covering correctness, coverage, consistency, exhaustiveness, and test-case coverage, and each reviewer traces the requirement through the codebase before it is allowed to read the plan. Then an author agent validates every blocker against real code. All of this is deterministic workflow code with severity gates, round caps, token reserves, and a formal dispute-escalation path, and when the disagreement is genuine that path hands it to the human rather than resolving it automatically.

Everything runs under budgets and stop conditions: spawn caps, a 3-strike rule keyed on root-cause tags rather than raw failure counts, broken-baseline hard stops, and an operator-owned-decision rule that stops the loop and hands back whenever a genuine product call comes up. Subagents never return prose. Every fan-out is schema-validated so that orchestration is always consuming typed data.

The far end now reaches production. I wrote the deploy stage by encoding what I had improvised by hand during one release night: a live-state pre-flight that caught a missing database default-privileges grant, a pipeline watch that turned a base-image CVE failure into a dependency-bump PR within minutes, and an unprompted filtered log watch while the team tested. The command reads live production state first and builds a go/no-go card from it. It computes the last green run from stage results, because a non-gating post-deploy test stage makes every run read as failed, lists the pending migrations in dependency order, checks the grants, and requires a rollback answer with a named database-restore owner before the card can go green. On a green card it queues the backend pipeline and then the frontend one, and a named human approves the pipeline gate every time, because the agent has no approve path at all. Every red stage is classified against a fixed playbook, and the ledger write lands before any action. Verification compares each workload's revision, image tag, traffic split, migration head, grants, and cache headers against expected text written into the release ledger ahead of time, so no row can pass by construction, and an unreadable check is recorded as blocked rather than failed, which means a tunnel outage can neither open the pointer bump nor print the rollback commands. Then a soak reads the error logs every ten minutes, subtracts a noise list and everything already seen, and turns every new signature into a row for a human to look at. The release ledger itself is a self-contained HTML page committed to git with append-only events and no server behind it, and the close-out tags the release on the ledger's own merge commit.

// pipeline — intake to production
  1. intake
    prototype · PRD · ticket · bug · transcript · prompt
  2. spec-decompose
    map the intent, freeze prototype anchors
    scope
  3. spec-reconcile
    diff against the shipped specs
    accept / reject
  4. spec-author
    mint permanent, immutable IDs
  5. spec-to-tickets
    vertical slices wired by blocked-by edges
    epic approval
  6. compound-test-cases
    cases and spec gaps derived before any code
    spec rulingsnew · 2026-09
  7. compound-plan
    executable plan
  8. compound-debate
    5-lens reviewer panel argues against the plan
  9. compound-master
    9-step spine: sync · ledger gate · work · verify · review · fix · iterate · learnings · dial
  10. compound-drain
    N parallel lanes over the ticket graph
    PR merge
  11. compound-deploy
    pre-flight card · launch · red-stage playbook · verify · soak · close-out
    GO card + pipeline approvalnew · 2026-09
  12. monitor
    filtered error-log soak while humans test; release ledger and tag

marks a human gate. The two outlined stages are the September 2026 additions: one pulls QA's test-case derivation left of the build, and the other carries a merged range through the production pipelines and the watch afterward. Small work skips the spec funnel and goes triage → test-cases → plan → debate → master → PR.

// outcome

The same engine serves supervised daily work, unattended epic drains, and now the production release, and what changes between them is the throttle rather than the guarantees. It fails closed. If a budget is exceeded or a schema does not validate, the workflow halts and escalates instead of quietly degrading.

[03] survey

##cross-team-prompt-pattern-survey

I surveyed another team's prompt suite, found five patterns that travel, adopted them into a different greenfield codebase, and wrote up the playbook.

Prompt engineering is converging across teams, and as far as I could tell nobody was harvesting the parts that travel. So I surveyed a brownfield team's suite, roughly 6,800 lines with 21 slash commands and 5 subagent templates, and I read it the way you would read another engineer's library, asking what is generalizable and what is load-bearing on their particular context.

Five patterns travelled: risk-driven mode selection, file-domain isolation between subagents, task-log files as durable phase contracts, active-fixer review agents, and frontmatter-registered subagents. I adopted all five into a different team's greenfield suite, and then I wrote a teammate-facing report so that other AI engineers could pilot the same patterns without having to redo the archaeology themselves.

// outcome

One engineer's archaeology lifted the capability of several teams. The skill underneath it is pattern extraction across codebases, and what it really takes is knowing which conventions survive translation into a new codebase and which ones were only scaffolding for the old one.

##contact

// get in touch
[loc]
Redondo Beach, CA

// open to roles and engagements building AI SDLC · enterprise engineering teams · AI labs