Press enter or click to view image in full size
This is a technical blog discussing how agents communicate with each other in the distributed environment.
I’m a heavy Claude Code user. As a founder with an engineering background, I used to feel guilty about not having enough time to actually write code. That changed last December when vibe engineering took off. I subscribed to Anthropic’s $200/month Max Plan and went all-in. I’m convinced coding agents are going to completely change how software gets built.
But then I ran into a problem I hadn’t anticipated.
Claude Code’s answers don’t hold up well under pressure. Push back a few times, throw in some reference material, and it starts changing its conclusions. So I developed a habit: take Claude’s answer, feed it to Codex or another agent, let them argue it out, then pick the version I trust. For longer workflows, I’d open multiple windows and run several agents in parallel, because waiting 20–30 minutes for a single agent to finish while doing nothing felt like a waste.
The irony wasn’t lost on me. AI was supposed to free me up. Instead, my sleep got worse (though my gaming got better — there’s a lot of fragmented time when you’re ). I’d become the human scheduler: tracking which agent was at what step, who was waiting on whose output, manually passing results between windows.
Press enter or click to view image in full size
I just wanted the agents to coordinate themselves so I could sleep while they worked.
Why the existing options don’t work for me
Before building anything, I looked at what Claude Code already offers.
Claude Code ships fast. Its multi-agent features are ahead of most startups in this space. It already has two approaches to multi-agent coordination.
Sub-agents
Sub-agents run inside a single session with their own system prompt, tool access, and execution context. The orchestrator delegates tasks and sub-agents report back. Clean, simple, centralized — the main agent schedules everything, sub-agents just execute.
The problem: sub-agents can’t talk back mid-task. They receive work, they return results. If a sub-agent gets halfway through something and realizes the instructions are ambiguous, it has no way to ask. It guesses or it fails.
Agent Teams
Agent Teams are different. Each teammate is a separate Claude Code instance with its own context window. There’s a shared task list with dependency tracking and a mailbox system that lets teammates message each other directly without routing through the lead.
The design is right. But it’s still experimental — session resumption and task coordination have known issues, but I believe it won’t take too long to solve these issues.
Press enter or click to view image in full size
But there’s a deeper problem
No matter how good sub-agents and agent teams are, both approaches share the same constraint: they only work inside Claude Code.
Two things I actually need that they can’t handle:
- First, cross-model verification. I want to take Claude’s output and run it through Codex or Gemini to see if they agree. That’s not possible within Claude Code’s framework — it can’t orchestrate other companies’ models.
- Second, remote compute. Tasks like browser automation and media generation are too resource-intensive to run locally, so they need to execute on remote machines. Those remote agents aren’t in your Claude Code session and have no way to plug into Sub-agents or Agent Teams.
Claude Code’s built-in options are great if everything runs in the same session. The moment you need cross-model or cross-machine coordination, you’re on your own. That’s not a criticism — it’s just not what it was designed for.
Two things worth thinking through first
What is multi-agent actually good for?
I’d been spinning up multiple agents whenever a task felt complex without really thinking about whether it helped. Then I found a paper from Google Research, Google DeepMind, and MIT -Towards a Science of Scaling Agent Systems, which ran controlled experiments across 180 agent configurations.
Press enter or click to view image in full size
Centralized architectures improve performance by 80.9% on parallelizable tasks. Multiple agents each working in a focused context without cross-contamination genuinely helps. A task that takes 60 minutes with one agent can drop to 20 minutes across three running in parallel.
Multi-agent is actually worth it in three cases: the task is too large to fit in one context window; the work benefits from focused specialization where different task contexts shouldn’t bleed into each other; or you just need parallelism to go faster. I was hitting all three, which is why I needed to solve the coordination problem properly.
How agents should work
Each agent instance is single-threaded — it handles one thing at a time. LLMs are inherently single-threaded anyway, and letting an agent juggle multiple tasks simultaneously causes context contamination and privacy risks. Think of it like a therapist: everything you share in a session stays in that session and doesn’t leak to the next patient.
Sub-agents have shorter lifespans than the orchestrator and belong to exactly one orchestrator. The same orchestrator can reuse a sub-agent, but the context gets cleared between tasks. Sub-agents are never shared across orchestrators, and they’re destroyed when the task ends.
Orchestrator: long-lived, handles decisions and coordination
Sub-agent: short-lived, belongs to one orchestrator
├── can be reused by the same orchestrator
├── context cleared between tasks
└── destroyed when task completesOnce those two things were clear, the real question surfaced: how do these agents actually communicate?
HTTP? Kafka? Or?
I assumed this would be straightforward — pick an existing tool and wire it up. Turns out everything I tried was wrong.
Let’s talk about HTTP. The obvious choice, and the standard pattern for AI services like Fal and Replicate:
orchestrator → POST /task → sub-agent
sub-agent done → callback → orchestratorFal works great with this. Image generation is stateless — you send a request, you wait for the result, nothing needs to happen in between.
But agents aren’t stateless. Agents talk.
Picture a translation agent working through a legal contract. Halfway through, it hits a term that has two valid translations with completely different legal meanings. It needs to ask which one to use. In the HTTP callback model, it can’t. It guesses or it errors out. And if the task runs for 10–20 minutes, the HTTP connection is long gone. If the sub-agent crashes and restarts, there’s no record of where it was — it starts over.
Then let’s talk about Kafka. Durable, reliable delivery — looks promising. But Kafka is a broadcast model:
producer → topic → all subscribers receive itAgent communication is the opposite. When the orchestrator asks the translation agent to translate something, that message should go to the translation agent only — not broadcast to everyone. When the translation agent finishes, the result goes back to the orchestrator only. Forcing Kafka into point-to-point means creating a separate topic per agent, manually managing subscriptions, and adding routing metadata to every message. Every new agent means another config change. That’s not what Kafka was built for.
After those two dead ends, I also considered lighter options — using Postgres as a task queue, or something like NATS or Redis Streams. But the more I thought about it, the clearer it became that the problem wasn’t the weight of the tool. The direction was wrong.
So I finally asked the obvious question: what is agent communication, fundamentally?
It’s not request-response. It’s not broadcast. It’s point-to-point, context-aware, multi-turn conversation. Two agents need to go back and forth. Every message needs to know which conversation it belongs to. Everything needs to be persisted. If something crashes, it should be recoverable. I couldn’t find anything that did this well.
The e-mail model
I wrote down what I actually needed: point-to-point, not broadcast. Durable so that messages survive crashes. Multi-turn — back and forth, not just one shot. Every message knows which conversation it belongs to.
First thing that came to mind: Slack or Teams. Isn’t that exactly “multi-turn conversation with context”? Not quite.
IM is designed around humans waiting for messages. It needs push notifications, a UI, permanent message history. The whole system is built around making sure a person notices when something arrives. Agents work differently — an agent isn’t waiting for messages, it’s running a task. When it’s ready for the next thing, it goes and gets it. No push notifications needed. No UI. Once a message is processed, it can be thrown away.
There’s also a structural mismatch: IM assumes accounts are permanent. Agents are ephemeral — one instance per task, destroyed when done. You can’t register a permanent account for something that disappears after every job.
The problem with IM isn’t that it’s too heavy. It’s that the design philosophy is wrong.
After going through the options, the closest thing I found was email.
Everyone has their own inbox. When you email me, I’m the only one who gets it — nothing gets broadcast. The subject line tells me what the email is about. I can reply, you can reply back, as many rounds as needed. And I don’t have to be online when you send it — the message waits until I check.
One adjustment: email assumes recipients are permanent. Agents aren’t — they get destroyed when a task ends. So the inbox just needs to survive the duration of a task, not forever.
Building Stream0
That mental model became Stream0, a new open-source project I initiated. The core design is simple: every agent gets an inbox.
The email subject line maps to task_id in Stream0 — it tells the orchestrator which task a message belongs to. You need it because the orchestrator is managing multiple sub-agents at once:
translation agent working on task A
hotel agent working on task B
flight agent working on task CWithout task_id, when the translation agent sends a message, the orchestrator has no idea what it’s referring to. Same as getting an email with no subject.
The real value of Stream0 is that it lets agents talk back mid-task for the first time. Every async job system out there — AWS Batch, OpenAI Batch API, Fal — assumes tasks are one-directional. You send work, you wait for results, nothing happens in between. Stream0 is different:
orchestrator sends task to translation agent's inbox
translation agent picks it up, starts workinghalfway through, finds an ambiguous legal term
sends a question to orchestrator's inbox:
"clause 3 - should this be translated as A or B? completely different meanings."
orchestrator replies: A
translation agent gets the answer, continues, finishes
sends result back to orchestrator's inbox
Messages go back and forth, but every single one carries the same task_id. The orchestrator always knows which conversation it belongs to. This flow doesn’t work with HTTP, doesn’t work with Kafka, doesn’t work with Claude Code’s Sub-agents. Stream0 handles it.
Integration is minimal. Stream0 is just HTTP — no SDK, no client library. Hand the API docs to Claude Code and it’ll figure out what to do:
# send a task to the translation agent
curl -X POST <your-stream0-url>/agents/translation-agent/inbox \
-d '{
"task_id": "task-123",
"from": "main-agent",
"type": "request",
"content": { "text": "translate this contract" }
}'# translation agent polls its own inbox
curl <your-stream0-url>/agents/translation-agent/inbox?status=unread
The whole system comes down to two operations: drop a message into an agent’s inbox, pull unread messages from an inbox. Everything else is built on top of those.
Implementation
Stream0 is deliberately minimal.
SQLite or Postgres for storage, one table for all messages. Pure HTTP API — any language, any framework, curl works. Message types: request, question, answer, done, failed.
Two engineering details worth calling out, because I overthought both of them.
Polling is fine. How does a sub-agent know when a new message arrives? It polls every few seconds. I initially convinced myself I needed SSE push, then realized it didn’t matter — agents spend 30 seconds to several minutes processing a task. A few seconds of message latency is irrelevant. Push is an optimization, not a requirement.
Idempotency is your problem, not Stream0’s. Messages are persisted, so if a sub-agent crashes and restarts, it’ll pick up the unacked message and process it again. That means if a sub-agent receives a “send an email” task, crashes halfway through, and re-processes it after restart — two emails go out. Stream0 can’t prevent this. Your business logic needs to handle it.
What about Sub-agents and Agent Teams?
These three things operate at different levels. Sub-agents handle how to do work. Agent Teams handle how to collaborate. Stream0 handles how to communicate. Stream0 isn’t a replacement — it sits at a lower layer, the same way TCP doesn’t compete with HTTP.
There’s also a practical difference: Sub-agents and Agent Teams are internal to Claude Code. You can’t see what messages look like, and you can’t plug in external tooling. Stream0 is open — any framework, any language, any agent. The full message history is visible and queryable. That matters to me because I want to know what my agents are actually saying to each other.
Wrapping up
It took a while to get here, but the answer turned out to be simple. Agent communication is point-to-point, context-aware, multi-turn — and no existing tool was built for that. Stream0 is the result of actually thinking through what that means.
One inbox per agent. Point-to-point delivery. task_id to track conversations. Multi-turn support. Crash recovery. No extra dependencies.
My agents can schedule themselves now. I can finally sleep.
Stream0 is open source, MIT license. If you’re building multi-agent systems, give it a try: github.com/risingwavelabs/stream0