Press enter or click to view image in full size
Here is a problem nobody warns you about: an AI agent running on a rented server in a data center cannot browse the web. Not “slowly.” At all. The same agent running on a Raspberry Pi in my living room works fine.
I found this out the expensive way, about halfway through building Helm. Here is what I learned.
What I was trying to build
Helm started as a security layer. The plan was straightforward:
- Run AI agents inside containers, in the cloud, always on.
- Put a checkpoint in front of them so they cannot take real actions unsupervised.
- Verify the risky stuff out of band, meaning through a separate channel the agent does not control.
Think of it like a bank teller and a vault. The agent is the teller who talks to you and figures out what you want. Something else entirely holds the keys.
That part worked. Almost nothing else did, and the failures were more interesting than the plan.
Roadblock one: your subscription is not for robots
The first wall had nothing to do with code. OpenAI and Anthropic both restrict flat rate subscription accounts when the usage does not look like a human sitting at a keyboard. An agent grinding away unattended at 3am is exactly the pattern they screen for, and reasonably so, since flat pricing assumes a human’s appetite.
I moved to MiniMax, which offers a good model and a subscription key you can hand to an agent without that constraint.
The catch is rate limits. Requests fail regularly, which means retry logic stops being a nice touch and becomes load bearing infrastructure. If you try this, plan for it up front.
Roadblock two: the web can tell where you live
This is the one that killed the original idea.
Every internet request carries a return address. Addresses that belong to data centers (Fly.io, AWS, and friends) are easy to spot, and a lot of websites simply refuse them. It is spam defense, and it works.
So my cloud agents could think, but they could not go outside. I gave one a headless browser, which is just a web browser with no screen attached, and the browser itself worked perfectly. It just could not get through any front doors.
Meanwhile the same agent on a Raspberry Pi at home, using a normal residential connection, sailed through. The takeaway, which I did not expect to write down:
A home internet connection is a feature, not a limitation. If your agent’s job is to go look things up, running it on your own hardware may beat running it in a data center.
What actually held up: keeping the AI out of the keys
Here is the design decision I would keep in any future version.
The Helm platform itself contains zero AI. Every decision it makes is a rule. Permissions, roles, allow lists. It can start agents, configure them, and shut them down, but it never reasons about anything.
Watch what happens when an agent wants to send an email:
- The agent writes a draft and asks the Helm backend to send it.
- Helm checks permissions and pings me on Telegram.
- I approve on my phone.
- Helm sends the email using credentials the agent never sees.
The AI proposed. It did not approve, and it did not execute.
That gap sounds pedantic and it is the entire safety story. The tempting shortcut, where the agent holds the credentials and “asks permission” as one more step in its own reasoning, looks identical from the outside and guarantees nothing. A model that can talk itself into a plan can talk itself past a checkpoint it also controls.
A nice side effect: the same approval queue handles logins. When a Google or Outlook connection expires, the reauth request lands in the same place. No new interface required.
The pivot: stop hosting agents, start serving them
Once cloud agents were clearly not paying for themselves, I pointed the whole thing at Claude instead.
By then Helm managed roughly two dozen connected services. Plugging those into Claude solved real annoyances immediately: multiple Gmail accounts side by side, services grouped per project instead of dumped into one pile.
Helm went from being a runtime to being a control plane for someone else’s runtime, and got more useful in the trade.
Memory, and the surprise it produced
Every agent I ran had the same gap: no structured, visible memory. Obsidian was my reference point, because you can open it and read what is in there.
So I built something similar inside Helm. Linked notes, indices, wiki style, but maintained entirely by agents rather than by me. Simple system. The value was consistency: same memory, same conventions, no matter which agent was driving.
Then the same trick solved a problem I had been ignoring.
Skills, meaning the instruction files that teach an agent to do a specific job, are miserable to distribute. There is no App Store for them. You email someone a file, they install it by hand, and every update repeats the process. Version tracking becomes folklore.
Serving skills through the same connection that already handled permissions turned the security layer into a distribution channel. Update once, everyone has it.
The onboarding trick
The last piece is my favorite, and I did not plan it.
Setup is now one paragraph of text plus one bootstrap skill:
- The user connects Helm in Claude.
- They paste a short prompt telling the agent to fetch the boot skill.
- The boot skill installs every other skill, wires up the connections, and walks the user through the configuration it cannot do alone.
The agent assembles itself and coaches you through the rest. For a non technical user, the barrier drops to “paste this.”
It is still too fiddly. The next job is having Helm run its own setup instead of narrating it.
The short version
I set out to build a hosting environment for AI agents. That is the part I no longer use.
What survived was three things: a memory the agents maintain, a permission layer that never reasons, and skills that distribute themselves. The agent runtime turned out to be replaceable. The scaffolding around it did not.
If you are running agents yourself: where did your bottleneck turn out to be? I would bet it was somewhere other than the model.