Consider an agent that calls the same tool several times with slightly different arguments. Half the calls fail. It rereads a file, follows two dead ends, and consumes far more tokens than usual. Eventually it produces an answer, and the user replies, “No, you’re looking at the wrong table.”
That session contains several clues about what needs fixing. But recording those clues does not automatically turn them into improvements to the agent’s memory, skills, or tools. The next session can repeat the same mistake.
We have been thinking about this while building self-evolving agents. An agent can already rewrite prompts, update memory, modify skills and tool definitions, and, in some environments, patch its own code. We want those abilities to grow out of the agent’s own experience: a correction from a user, a failed tool call, or a successful task that took far too much work. The challenge is connecting that experience to changes the agent can use in later tasks.
We use three ideas to organize that work: pain, reflection, and sleep. Pain records negative feedback. Reflection turns it into a testable explanation. Sleep gives the system time outside user tasks to examine experience and try changes. Evaluation determines which changes we keep.
Press enter or click to view image in full size
Pain gives reflection a concrete starting point. Sleep looks across sessions for further opportunities. Both contribute to changes used in later tasks.
Pain: capturing what went wrong
We call the signals in that opening example pain events.
Some signals come directly from users. “That’s wrong” and “I already told you this” are easy to recognize. Others need more interpretation: a user repeatedly corrects the same mistake, regenerates an answer, undoes an action, takes over a workflow, or abandons the task. None of these proves that the agent is broken. Across many sessions, though, they can help identify a recurring problem.
The environment supplies evidence too. Invalid API arguments, repeated tool failures, timeouts, permission errors, failed commands, rollbacks, and retry storms can reveal a poor strategy. One timeout may be noise. The same invalid API call in twenty sessions is a reason to inspect the skill or tool definition.
The category we find most interesting is behavior during tasks that succeed. Suppose a simple request consumes 80,000 tokens, the same file is opened six times, or the agent alternates between two hypotheses without making progress. A correct final answer can hide a costly path to it. We think of this as behavioral pain, and we want the system to notice it before the user has to.
Recording these events gives reflection evidence to work with. Frequency, severity, and impact help distinguish a recurring weakness from an isolated failure. A repeated correction across hundreds of sessions, for example, gives the system a stronger basis for investigation than one timeout.
Reflection: producing a testable change
Pain identifies a problem without necessarily explaining its cause. Reflection has to do that next piece of work. Asking a model to critique its last answer often produces advice such as “I should have been more careful.” That gives us little to test.
Suppose an agent repeatedly fails while paginating through an API. Reflection might find that the failed calls use an obsolete page argument, trace it to an outdated example in a skill, and propose a cursor-based example. We now have evidence, a suspected cause, a file to change, and an expected outcome that we can check against the failing tasks.
That is the standard we want from reflection: enough specificity to produce a diff and evaluate it.
Choosing the affected layer is part of the diagnosis. A forgotten user preference may belong in memory. A misunderstood workflow may need a better skill. A tool interface that encourages invalid calls may need a schema change. Repeated loops can point to the runtime or control logic.
The possible changes span memory, prompts, skills, tools, plugins, control loops, and application code. We think of this as the agent’s evolution surface. Cost and risk usually grow as changes reach deeper into the stack, so the aim is a small change at the layer responsible for the failure.
Sleep: learning across sessions
Pain and reflection cannot catch every problem. Some inefficiencies produce no obvious negative feedback, and deeper patterns may only become visible across many sessions. Sleep gives the agent a regular opportunity to look across sessions, including those that appeared successful, and discover opportunities for improvement that individual pain events and reflections missed.
We call the separate period for examining experience and testing changes sleep. During this offline consolidation period, the agent can review recent sessions, group similar failures, propose changes, run experiments, and decide what belongs in the next version.
For example, a sleep cycle might inspect several thousand sessions and find a database error in 8% of analytics tasks. Looking at those sessions together could reveal one ambiguous skill instruction, or three unrelated problems with similar symptoms. The broader view helps the system distinguish them.
This gives sleep a broader view than reflection on a particular complaint. It can revisit how tasks were completed, connect observations across sessions, and discover improvements even where nobody reported a failure. Running that analysis outside the current task gives the system room to work through what it finds.
A proposed change still needs a check against the current agent and existing regression cases. The component proposing it should have limited authority over its evaluation, and code changes belong in an isolated environment with review and rollback. These checks let the learning process produce changes we can inspect and use.
Building the loop in MoziForge
We are building MoziForge, an open-source framework for developing self-evolving agents on DeepSeek Harness. It brings session evidence, feedback collection, reflection, and training into one workflow, with isolated evaluation and human-reviewed integration. The repository includes the plugins and documentation for running the system and following a change through that workflow.
We chose Harness because its plugin architecture gives us access to the parts an evolution system needs to observe and change: sessions, tool execution, and the agent loop itself.
A plugin is a module loaded into the running application. It can provide a service, register tools, or respond to runtime events. Harness uses Cordis to compose these modules through configuration. Its model adapters, tool registry, session log, and agent loop are themselves plugins, so the same extension mechanism reaches much of the agent’s execution environment.
That matters for MoziForge in two ways. Session and tool events give us evidence about what the agent actually did. Configurable components give us places to apply improvements, including changes to tools and runtime behavior. We can add the feedback and learning workflow through public Harness services while reusing its session handling, model providers, and execution facilities.
The agent-pain-plugin records feedback and execution signals as durable evidence. The reflect-loop-plugin groups that feedback and delivers reflection work to Trainer. The sleep-loop-plugin uses session summaries to deliver broader analysis work, including opportunities found in sessions that appeared successful. Trainer develops the resulting plans and changes; Agent Test provides isolated evaluation, followed by reviewed integration.
Press enter or click to view image in full size
Harness supplies the execution environment; MoziForge plugins connect session evidence and feedback to improvement work.
From reflection to new capabilities
The report below shows MoziForge examining its own feedback collection. It distinguishes problems that need changes from expected tool behavior, then connects the findings to a plan for review. One finding concerns the collector counting Trainer’s own reflection sessions as pain; another concerns how cached tokens contribute to a high-usage signal.
Press enter or click to view image in full size
MoziForge can also respond to a missing capability by building a plugin. In a recent session, We asked it to connect to my iCloud Calendar so it could analyze my calendar data. It developed the integration, created a GitHub repository, and published the source as calendar-plugins.
The iCloud plugin connects through CalDAV and exposes tools for listing calendars, reading events, and creating an event. It expands recurring events into individual occurrences so the agent can work with the events in a requested time window.
Press enter or click to view image in full size
This is why the plugin architecture matters to me. A request that needs a new connection can lead to a reusable addition to the agent’s environment. Prompt and skill changes help it use the capabilities it has; a new plugin gives it access to something it could not reach before. The calendar request produced code that can be inspected, used again, and developed further.
Where we hope this leads
We want MoziForge to become more adaptable as it works with a user. That means learning how the user wants things done and developing the capabilities needed to do them. Over time, We hope it can improve nearly every part of how it works, from choosing a tool to organizing a team of agents. Plugins give that ambition a practical form: an agent can add a connection to an external service and use what it learns there in subsequent work.
The longer-term idea we keep coming back to is an agent company. Given a high-level goal, MoziForge could organize a team that builds missing software, carries out the work, and improves its methods along the way. Like a small technology company, it would need to discover what it lacks and decide how to acquire or build it.
For a project that needs customers, that could include finding leads and reaching out by email. For research, it could gather current news and follow new information as the work develops. It could search for existing plugins, adapt them, and contribute a pull request when an improvement would help other users. We imagine the team interacting with people and services as part of pursuing the goal, with those interactions giving it further experience to learn from.
That is the direction we want to explore with MoziForge. Pain preserves the difficulties encountered during work, reflection develops possible changes, and sleep gives the system a broader view of its experience. The calendar plugin is a small example of the capability growth we have in mind. We would like to give the agent a goal, work through the first attempt together, and find that the next attempt starts with the tools and lessons we gained from the last one.