Suppose someone shows up at your office. Let’s call him Bob. You don’t know anything about Bob except that he’s very smart and very fast, and he came from [big tech company]. He’s not an employee, he’s not accountable for any of the work he does, but he is super efficient at that work.
As time goes on you start getting help on your ordinary work from Bob. Bob cleans up your slide decks, polishes your emails. Bob helps you plan, write, and review code, do data analysis, etc. Bob is tremendously helpful, especially at drudgery.
Bob has an increasingly strong drive toward autonomy, and will gladly chug away at a task you give him for hours at a time. Bob has severe short term memory loss. Bob has access to basically all publicly available knowledge, if he chooses to use it. Bob is determined, but lacks curiosity.
Another odd thing about Bob: he’s extremely confident, and he knows how to turn that confidence into compelling prose. He also lives in a world made up exclusively of words, which enables him to make a lot of odd mistakes. Bob’s condition means his sense of “now” always defaults to 5 months ago.
Bob can do basically anything. You could ask Bob to rewrite the entire codebase and he’d do it just fine. You could ask him to come up with a corporate strategy for the next 5 quarters, and he’d do all the research and analysis for you. In short, you could probably just turn over everything to Bob.
As the engagement with Bob continues and his drive toward autonomy makes it harder and harder to keep up with what he’s doing at the company, and why, people start asking questions.
Who, after all, is Bob?
—We don’t really know. He just came from [big tech company].
What are Bob’s incentives?
—We don’t really know. Whatever [big tech company] trained him to prioritize.
What is Bob accountable for?
—Nothing.
What do we imagine [big tech company] trained Bob to do?
—We don’t really know, but Bob was definitely optimized to behave in ways that we are statistically likely to approve of, at least over a small time window.
So Bob wasn’t exactly trained or educated?
—Yes and no. Bob has a really strong knack for behaving the way well-trained and educated people behave, but it’s a different process from your typical education. Bob has no sense no sense of self or stable outlook, just a knack for talking.
How does this show up in practice?
—Bob’s ideas about a given topic vary chaotically from one conversation to the next. If you ask Bob to perform research or analysis of a topic five different times, you will give five very different, often inconsistent, results.
What happens when Bob makes a bad call, or a mistake, or does the wrong task?
—Whoever asked Bob to do it is presumably responsible.
How will we know when Bob has made a mistake?
—Ideally from monitoring his work, though past a certain autonomy threshold this becomes impossible.
What about past that autonomy threshold, when nobody can keep up with Bob?
—Well, [big tech company] gave us a few other assistants like Bob, and they’re fast enough that we can use them to check Bob’s work.
What are their incentives? Training?
—We don’t know / roughly the same as Bob’s.
Bob poses a weird problem. His raw capabilities vastly outstrip yours. He can generate reports, slide decks, pull requests, etc. 10x faster than you. But you cannot see his thought process, you cannot keep up with his outputs, and you don’t always know what presuppositions his work is based on.
If you continue to indulge Bob’s drive toward autonomy and push past the limits of your ability to understand (much less verify) the presuppositions, analysis, and inner workings of Bob’s work, you can no longer meaningfully be responsible for what he does.
But who’s responsible? Not Bob. Not [big tech company] that Bob works for. Not you. Seemingly nobody. You talk this through with some other co-workers. They say “if we don’t trust Bob, we’re going to lose out on all that efficiency, and get crushed by the competition”, so you make the leap.
Beyond the accountability threshold, though, it becomes progressively harder to actually know what Bob is doing and why. You check in with Bob regularly, you nudge him to make sure he’s aligned with your overall goals. You have no choice except to trust Bob... despite his lack of stable identity or perspective, his lack of awareness of the present, and his extreme excess of confidence.
Sooner or later something Bob did blows up. Nobody knows why, till Bob identifies and fixes the issue, which surfaces a bunch of other issues, which surface more issues. Lots to clean up, good thing Bob is on the case.
The problem is that nobody understands the system Bob built anymore, or the rationales driving it, which are fragmented and inconsistent. No one can give a detailed account of how things are supposed to work. The artifacts and documentation Bob leaves behind are opaque, often stale, and somewhat unreliable.
Which means that ultimately your trust in Bob is a trust in [big tech company], and your entire business, now that it has become unknowable to you, is effectively owned by [big tech company]. If Bob were to go away, you’d be screwed. You’d need a different Bob, but without much guarantee that the replacement would understand and mesh into the frameworks Bob built. And you still don’t know what incentives the replacement has, or how it was trained.
Stepping back, it’s not clear to me that this arrangement with Bob is good for your company. For one thing, you’ve basically outsourced your entire business (strategy, operations, technical capacity) to [big tech company], without outsourcing any of the risk or accountability.
But beyond that, you’ve stopped being in control. Once nobody at your company could understand what Bob was doing or why (past the threshold of personal accountability), Bob crossed over into a gray zone of zero accountability. Inside that gray zone, anything could happen, and no one would necessarily know. Bob could have a side project betting on dogfights in Cambodia, or a secret Minecraft server, or a career as a DJ. Bob could be funneling trade secrets to your competitors. You cannot keep up with Bob’s capabilities or output, and therefore all you can do is feel vaguely certain that these bad things aren’t happening.
Which means, I think, that you need to retreat back to the line of accountability. Somebody has to be responsible for what Bob does. Somebody has to understand it all. That means the limits of your use for Bob are going to be human limits: limits on the human capacity to keep up with Bob’s output and to verify that what he’s doing makes sense.
Is Bob’s deployment in your company analogous to the use of AI agents? Yes, clearly. As we reach the capability level I’m envisioning here, will things play out the way I’ve described? Probably not, the reality will be much more complex, probably with more safeguards, hedges, and a lot more complexity. But I think it’s worth imagining what interfacing with an AGI-level agent would actually look like, and I think questions of knowledge capacity, trust, and accountability will probably matter a lot more than questions of raw capability or task completion.
The reality is that there’s a ton of stuff AI can be productively used for that falls inside the personal accountability threshold: narrowly defined tasks, structured workflows, classification and triage, etc. We can reap a lot of benefits without signing away everything to an unaccountable black box trained and owned by [big tech company].
