In a recent LinkedIn post, Linas Beliūnas highlights a critical bottleneck in the development of artificial intelligence agents: the tendency to allow them to self-evaluate. Beliūnas shares a profound insight, attributed to an unnamed Anthropic engineer, that strikes at the heart of AI improvement.
According to Beliūnas, the path to truly self-improving AI agents involves a fundamental shift in their evaluation process.
“Your AI agents will never get better until you stop letting them grade their own homework. Once you give it external feedback, the system finally starts improving itself.”
This statement, as shared by Beliūnas, underscores a core principle that many in the AI development field are grappling with. The idea of an AI agent being its own judge is presented as a significant impediment to genuine progress.
Separating the Worker from the Judge
Linas Beliūnas elaborates on the practical implications of this feedback loop, referencing a 15-minute talk by an Anthropic engineer. This talk, according to Beliūnas, outlines a concrete system designed to foster self-improvement in AI agents. The central tenet of this system is the separation of the AI’s operational function (the “worker”) from its evaluative function (the “judge”).
Beliūnas points to the engineer’s proposed methodology, which involves:
- Introducing genuine external feedback into the system.
- Scaling this process through the use of sandboxes, memory mechanisms, and multiplayer harnesses.
This structured approach, as presented through Beliūnas’s post, suggests that true AI advancement hinges on objective, external validation rather than internal, potentially biased, self-assessment.
Building an Agentic OS from Scratch
Further emphasizing the practical applicability of these concepts, Linas Beliūnas shares a resource for developers looking to implement these principles. He notes the availability of a guide that enables the creation of an “Agentic OS” from the ground up, utilizing Claude Fable 5.
“In this 15-minute talk, an Anthropic engineer shows the exact system that actually makes agents self-improve: separate the worker from the judge, add real external feedback, then scale it with sandboxes, memory, and multiplayer harnesses.”
This accessibility to a detailed system, directly from the creators of advanced AI models like Claude, is highlighted by Beliūnas as “pure signal.” It indicates a move towards more robust and reliable AI development practices.
The Importance of External Validation
Linas Beliūnas’s sharing of this information serves to highlight a crucial paradigm shift in AI development. The traditional model of AI learning, which might rely heavily on internal metrics or self-correction, is being challenged by a more effective approach that prioritizes external, objective feedback.
As Beliūnas conveys, the insights shared by the Anthropic engineer provide a clear roadmap. By distinguishing between the AI’s task execution and its performance evaluation, and by incorporating real-world data and feedback, developers can create agents that possess a genuine capacity for self-improvement. This distinction is vital for moving beyond incremental gains to achieve significant leaps in AI capabilities.
📝 About This Content
This article is based on insights shared by Linas Beliūnas on LinkedIn.
📅 Originally posted on July 24, 2026 | View original post on LinkedIn →