Press enter or click to view image in full size
My team hasn’t written a line of code by hand all year, across more than 2,500 commits. Still our process felt too manual. In 2026, the dream (expectation?) is to simply type “build X” and have an improvement arrive on prod. So this summer I decided to do my best to make that happen.
Spending about a month building out Orca (a product engineering orchestrator customized to our needs), I overdid it and learned — at least today, for products with significant complexity — there is such a thing as too much automation. A dialectic process led to a middle ground where, at the moment, things feel perfectly tuned. A small kernel of human engagement keeps the process on rails, ensures high quality, respects security, and scales to tasks of significant complexity.
Most important is the simplicity of Orca’s human-facing interface. The team is left with a crisp, airy, collaborative development experience that requires from humans only the essentials. The attention/time required to get things done appears to hit a lower bound.
My experience led me to a surprising observation. For nuanced product engineering over the past year, input seems to be remaining constant while output increases dramatically. If the kernel remains as LLMs further improve, the future will not be fully automated.
Nuanced needs called for a custom solution
The famous stories of automation — porting Bun to Rust, the Ralph loop— all have something in common. They describe situations where the end goal can be clearly defined. We found ourselves with a much messier task shape. Our product, Moonchild, is an AI-native design system agent and UI/UX design bot. Our constraints resemble much of what I’ve seen in industry:
- Real infrastructure: our GCP topology has gotten quite nuanced, requiring a VPN to run fully.
- Security: we handle sensitive customer data, and our system is riddled with keys tied to billing accounts.
- A high bar for UI: our user base of pros in the field is highly discriminating. For many of our features there is no prior art.
- Agentic: product flows are fundamentally non-deterministic. Despite evals, human judgment is still needed to validate functionality.
- Feature sophistication: there are no easy waterfalls — most meaningful tasks bounce between technical design, UX, and product as they are refined and resolved.
- Cost: we can’t afford to tokenmax and burn compute. As an early-stage company, we need to be efficient with our time too.
Our dev rigs were already well harnessed. What we were missing was the orchestration layer that coordinates and takes tasks end-to-end (not just building but scoping, architecting, designing, and validating). Existing orchestration solutions did not meet our needs. OpenClaw was too difficult to control and perhaps even more difficult to trust. Cursor Cloud could not handle our stack without gymnastics and, like Claude Code, required significant additional work to truly be an orchestrator.
The solution was Orca, a custom orchestrator running as a Node/Express server on a Mac Mini with Claude Code sessions doing the heavy lifting. Orca is controlled entirely via Slack, each issue is a Slack thread.
In code, I define exactly the bot’s control surface and its interactions with our setup (granular security, precise cost control and attribution, integrations with things like Sentry, GitHub and Linear). This approach is in line with current wisdom that repeated processes should be moved from inference to hard code.
When there are multiple north stars, full automation can churn badly
As I pushed Orca to its limits, I discovered something that — at the moment — feels fundamental. A well-defined kernel of human agency (made available via a clear-cut chat UI) makes everything better. The agent spins less, the right thing gets done, the right decisions get made, tokens are used efficiently, and peace of mind ensues.
I learned this the hard way, by building a fully automated agent. This system worked for small tasks. Bug fixes and small improvements progressed nicely. A small amount of conversation was typically needed to guide the fix, after which the correct PR would be produced on either the first build-out or after one or two loops of validation.
But things really fell apart for more complex tasks. For example, adding tabs to our chat system. Moonchild has long had multiple tabs per scene, accessible through a history menu. We wanted to add tabs, to make it easier to swap chats and make the underlying data model more clear to our users. This is exactly the kind of task where one should be able to simply type “add tabs to the chat” and get the right result. The best solution shouldn’t require any particular judgment, it’s an industry-standard move.
Press enter or click to view image in full size
Working from a simple scope, our fully automated agent made quite a mess of things. The experience was like one of those ChatGPT sessions that keeps flip-flopping between different recommendations, however with much more weight. Each turn cost many minutes and many tokens. Also, the residue from past turns ended up clogging the works and bringing the whole system to a halt. It was like our agent got lost in a hall of mirrors without end.
The anti-pattern was choices in one domain mucking up choices elsewhere. For example, consider the inclusion of an empty state in the chat panel. In the initial UX workup, this was set as a goal: if all chat tabs are closed, an empty state will be seen. As the agent went to architecting, supporting this UX necessitated deep structural changes — including a fundamental alteration of the poll cache mechanism we use to synchronize UI state across multiple browsers.
Press enter or click to view image in full size
The agent eventually figured out the empty state was a bad move (instead, users should not be allowed to close the last tab). However as it back-tracked, detritus of the prior folly clouded its context. And the empty state was just one of several concerns that clogged the system up in this way. Attempts by the agent to solicit human input — or the human to intervene — seemed to make things worse. The result was a mess that left us asking “is this actually faster?” All this for a task that, let’s be real, is not that big.
Multiple oracles make real-world, nuanced product improvements resistant to naïve automation
I diagnosed the underlying issues as twofold:
- The lack of a clear oracle. Long-running highly orchestrated tasks succeed when there is a single clear definition of success. Non-trivial features have to be optimized along at least three axes: scope, technical architecture, and UI/UX. The LLM churned because it was alternating between goals. Complexity compounded.
- Uncoordinated, overlapping capabilities: Our agent had plenty of tools, but was not able to use them together effectively. Almost every ticket turned into a mishmash where prior half-decisions stacked. Context got messy, quickly.
While the second point is a matter of proper agentic design, the first point proved more sticky and I believe fundamentally limits the automation of our work at this moment in time. It’s difficult to encode what we mean by quality in advance, even for straightforward tasks like adding tabs to Moonchild’s chat: base models do their best but to no avail. The actual choices we make are so often situational. We aim to build standard product, nothing eccentric, obeying the principle of least surprise both in technical and interface design. Yet time and time again, the agent misses. This can be seen even in smaller-scale, non-orchestrated, harness-led efforts. While I do believe that in the future this better oracle could be developed, it doesn’t seem ready yet.
I made things better by solving the oracle problem, which I’ll detail below. As for the agentic design problem (issue #2), the solution involved delicate management of sessions, Git worktrees, branches and context in the underlying Claude Code workers. Instead of the typical Claude skills, I ended up with a higher-level control mechanism that lives on the orchestration level above Claude Code itself. These technical decisions followed on from the approach to issue #1.
Modes not only made Orca run smoothly, but left us with a delightful developer experience
Solving the oracle problem, like any system design challenge, depended on finding the right abstraction. For us this was a set of modes that completely spans the task space with minimal overlap.
Press enter or click to view image in full size
These modes (each essentially its own agent) not only lead to increased token and time-efficiency, but also provide granular control over security (each mode has its own access profile — some have limited credentials, but none of them connect to GCP).
Key was to make each mode have a specific set of deliverable types (such as a technical design). These artifacts, ideal for human review and discussion, then feed the next mode allowing it to act with increased efficiency.
Press enter or click to view image in full size
I found these modes by doing a task analysis of the errant attempts at full automation to find the natural clusters of concern. It was at first surprising to see the modes so closely align with classical product-team task division, but then of course it made so much sense. Human minds have arrived at this structure after years — no, decades — of process work aiming to separate the phases of work cleanly. Since LLM intelligence is human intelligence materialized in silicon, it’s not surprising to see it excel with these groupings.
Isn’t LLM excellence supposed to be so much more… automated? I dove into Claude Code creator Boris Cherny’s process and found that, from a certain angle, his agents are exactly like Orca’s modes.
Working with Orca, I find the traditional product waterfall (scope → UX → architect → build) is rarely the best way to get things done. Many tasks start out with architecture before heading to scoping. Or sometimes the scope is so minimal it can be skipped and we can go right to the UX phase. For bugs and customer requests, investigate often starts the show after which UX may not be needed. Some tasks end up bifurcating or ping-pong between multiple steps before they are ready for the build mode. The lightweight mode mechanism makes this natural and efficient, here the deliverable system really shines.
The one constant is the simplicity of the build phase, and how routinely it delivers shippable PRs in a single shot. I am surprised to see the disappearance of the plan artifact seen in many agentic coding workflows. As someone who’s been doing plan-based development since before Cursor had a “plan” mode, I was shocked to see my workhorse disappear. Turns out, if the scope and UX and architecture are well elucidated, and you’re building with Opus 5 or stronger, the plan becomes implicit and need not be externalized.
Each mode has not only a demeanor, but a particular thing it’s optimizing and tools/deliverables to match. This alignment leads to a terrific efficiency in work. Seeing the garbled mess of our early orchestrator get replaced by clean modes, I knew we had found the right abstraction. The lengthy, circuitous discussions have been replaced with a few focused chat turns. The overall feeling is of providing the minimal amount of human input, with crystal-clear artifacts (terse markdown, preview builds, videos) available to make it easy to say the right thing next.
Surprisingly, in this new system, even small bug fixes are much cleaner than they were before (fewer conversational turns, more concise responses, cleaner code updates). The mode architecture has numerous advantages which are beyond the scope of this article (granular control of security and model usage, detailed cost management, and the ability to maximally leverage LLMs — each mode is an agent in the classical sense).
The net feeling is pleasant, perhaps because as primates it’s natural to work in groups. Since each mode is a persona, my brain can operate with the system in a way that it’s been selected for (in the Darwinian sense).
Is the kernel essential?
The evolution of Orca surprised me — instead of full automation, a kernel of human interaction (the mode system) remained. I’ve come to understand this as being due to accumulation of error. If each step of a multi-step process is 80% correct, quality drops quickly. For a five-step process, 0.8⁵ is ~33%, meaning the end result is 2/3rds wrong! (I am not the first to make this observation.)
And 80% seems fair, that matches my experience using dev harnesses to execute well-scoped tasks, using well-informed chatbots to produce scope, and well-powered design tools to produce UX solutions. As bots have evolved, they handle larger tasks, but they continue to have a tendency to bind early, drive into suboptimal parts of the solution space, and create a mess. (I must add that part of what’s nice about Orca is the bringing together of all these prompts into one single Slack thread. There is only one place to go, and all the phases of work can share context).
When there’s a clear, single oracle, iteration (long-running LLM sessions) can take that 80% up to 99%. But in our “real product” world (I don’t know what else to call our needs other than “real” lol) the oracles are nearly impossible to define, and always evolving. This provides for an interesting take on Moore’s law for AI: METR’s observation that LLM coding task length doubles every 7 months. This doubling rate is for benchmark tasks with clear oracles. Real work is more nuanced.
That said, I’m loath to put a stake down and claim that the kernel that works for us right now, this summer, is here to stay. It’s easy to imagine a more powerful LLM automating our modes in 6 months. My bold hypothesis is that — even if an LLM were to automate a bunch of what we’re doing manually — we wouldn’t arrive at full automation. Because a new kernel would then emerge.
Part of my support for this hypothesis is historical. Over the past 12–18 months of purely prompt-based development, my work hasn’t reduced. Instead, the amount of work (my inputs to the chatbots) has remained constant while output has grown considerably. And some of the leaders of orchestrated development, such as Boris Cherny, have a kernel too.
Instead of long-running turns that last all night, most turns are in the 3–5 minute range, maxing out around 15 minutes. Perhaps, instead of work being automated, work will remain as-is while output scales. This reframing of automation cuts against the current narrative that longer task lengths will dominate as LLMs evolve. By reducing work to the minimal surface needed for human cognition and interaction, I managed to greatly increase productivity while eliminating virtually all the mundanity of building a sophisticated, nuanced product. If you’d like to learn more about this topic, I can recommend this paper and this report.
If you’re working on a similar orchestrator, I’d love to hear from you. The process has been fascinating and I have another hypothesis that every team needs its own orchestrator. They are quick to build these days and every team seems to have its own needs. But of course I don’t have the experience to back that up (…yet?). Don’t be afraid to DM or add a comment. Also happy to share more technical details for those who are working on similar systems.