How I build software with AI agents
I don't only teach agentic engineering — it's how I ship. On a typical day I run nine or more agent sessions at once, mostly in Claude Code, working through a governed loop with GitHub as the single source of truth. The method I bring to a team is the one I use myself.
The loop, step by step
- Issue. Every unit of work starts as a GitHub issue. A Planner agent drafts the issues and I approve them before any work starts; approved issues are the only work queue.
- Claim. An agent claims an issue by creating its branch. The branch ref is an atomic claim — only one agent can win it — and labels track each issue's state.
- Branch. The agent works on that branch and nowhere else.
- Pull request. The result is an ordinary pull request with one concern, checked against the issue it claims to implement.
- Independent review. A different agent instance reviews it. No agent approves its own work. After two rounds of requested changes without converging, it escalates to me.
- Green CI. A red build goes to a Fixer agent, which diagnoses it and pushes a fix; repeated failures escalate.
- Merge. An approved pull request with green CI merges automatically — except on human-only paths.
Changes flow from integration to staging to production. Promotion to production is automated once staging checks are green.
Seven roles
- Planner — reads the roadmap and open issues, and files well-scoped tasks with clear acceptance criteria, so there is always ready work.
- Worker — claims exactly one task, implements it end to end and opens one pull request. If the task turns out bigger than expected, it ships the smallest useful slice and files a follow-up.
- Reviewer — a different instance from the one that wrote the code checks the diff against the issue and the project's conventions, then approves or requests changes.
- Tester — exercises the change the way a user or CI would, and reports what actually broke.
- Fixer — drives failing pull requests back to green, so Workers aren't tied up babysitting red builds.
- Janitor — a small, deterministic process with no model calls that releases stale claims and nudges quiet pull requests, so the loop can't silently stall.
- Orchestrator — supervises the fleet, keeps the right mix of roles running, and steps in where a single agent shouldn't decide alone.
What stays mine
- Secrets and human-only paths. Agents never see credentials, and permission deny-lists block them from payment and credential code.
- Approving the specs. The Planner drafts every issue; none enters the queue until I've approved it.
- Escalations and final calls. When the loop can't settle something, I do.
UTON OS: when the loop scaled
UTON OS is the booking platform, with payments, for UTON's recording studios — built for a paying client through this loop.
Its hardest problem wasn't a feature — it was scale. As the number of concurrent agents grew, two things broke. CI and review couldn't keep up with the volume of pull requests, and cost and provider rate limits climbed with concurrency.
I solved it by capping concurrency to what CI and review could absorb, and by making CI faster and tiered: fast required checks split from slower ones, with cached builds.
The lesson for any team: CI is the throttle. Agent throughput is capped by CI and review capacity. Invest there before adding agents.
Tools I build for agents
- Codewright — a risk-classified knowledge base of terminal recipes, served to agents over an MCP server and a Claude Code plugin.
- RallyPoint — my published, MIT-licensed protocol for multi-agent coordination. It shows how I think; each team gets a setup that fits it.