Live data from Hacker News

HyperAgents: Self-referential self-improving agents

github.com

21–30 of 117 posts

Re: HyperAgents: Self-referential self-improving agents

#21

The paper is here - https://arxiv.org/pdf/2603.19461 This, IMO is the biggest insight into where we're at and where we're going: > Because both evaluation and self-modification are coding tasks, gains in coding ability can translate into gains in self-improvement ability. There's a thing that I've noticed early into LLMs: once they unlock one capability, you can use that capability to compose stuff and improve on oth…

because submarine piloting is a going-under-water activity, improvements in holding one's breath can lead to faster submersibles.

Re: HyperAgents: Self-referential self-improving agents

#23

The paper is here - https://arxiv.org/pdf/2603.19461 This, IMO is the biggest insight into where we're at and where we're going: > Because both evaluation and self-modification are coding tasks, gains in coding ability can translate into gains in self-improvement ability. There's a thing that I've noticed early into LLMs: once they unlock one capability, you can use that capability to compose stuff and improve on oth…

I disagree that evaluation is always a coding task. Evaluation is scrutiny for the person who wants the thing. It’s subjective . So, unless you’re evaluating something purely objective, such as an algorithm, I don’t see how a self contained, self “improving “ agent accomplishes the subjectivity constraint - as by design you are leaving out the subject.

Sure. There will always be subjective tasks where the person who asks for something needs to give feedback. But even there we could come up with ways to make it easier / faster / better ux. (one example I saw my frontend colleagues do is use a fast model to create 9 versions of a component, in a grid. And they "at a glance" decide which one is "better", and use that going forwards).

OTOH, there's loads you can do for evaluation before a human even sees the artifact. Things like does the site load, does it behave the same, did anything major change on the happy path, etc etc. There's a recent-ish paper where instead of classic "LLM as a judge" they used LLMs to come up with rubrics, and other instances check original prompt + rubrics on a binary scale. Saw improvements in a lot of evaluations.

Then there's "evaluate by having an agent do it" for any documentation tracking. Say you have a project, you implement a feature, and document the changes. Then you can have an agent take that documentation and "try it out". Should give you much faster feedback loops.

Re: HyperAgents: Self-referential self-improving agents

#24

Pi is self modifying, self aware. https://lucumr.pocoo.org/2026/1/31/pi/ But this idea of having a task agent & meta agent maybe has wings. Neat submission.

What are the differences wrt Recursive Language Models

Completely unrelated. Recursive Language Models are just "what if we replaced putting all the long text into the context window with a REPL which lets you read parts of the context through tool calls and launch partitioned subagents", ie divide-and-conquer applied to attention space.

Re: HyperAgents: Self-referential self-improving agents

#25

No matter how far we go, we end up with generation / discrimination architecture. Its is the core of any and all learning/exellency; exposure to chaotic perturbations allow selection of solutions that are then generalized to further, ever more straining problems; producing increasingly applicable solutions. This is the core of evolution, and is actually derivable from just a single rule.

I don't think generation/discrimination is fundamental. A more general framing is evolutionary epistemology (Donald T. Campbell, 1974, essay found in "The Philosophy of Karl Popper"), which holds that knowledge emerges through variation and selective retention. As Karl Popper put it, "We choose the theory which best holds its own in competition with other theories; the one which, by natural selection, proves itself the fittest to survive."

On this view, learning in general operates via selection under uncertainty. This is less visible in individual cognition, where we tend to over-attribute agency, but it is explicit in science: hypotheses are proposed, subjected to tests, and selectively retained, precisely because the future cannot be deduced from the present.

In that sense, generation/discrimination is a particular implementation of this broader principle (a way of instantiating variation and selection) not the primitive itself.

Re: HyperAgents: Self-referential self-improving agents

#26

The paper is here - https://arxiv.org/pdf/2603.19461 This, IMO is the biggest insight into where we're at and where we're going: > Because both evaluation and self-modification are coding tasks, gains in coding ability can translate into gains in self-improvement ability. There's a thing that I've noticed early into LLMs: once they unlock one capability, you can use that capability to compose stuff and improve on oth…

The whole theme of llm dev to date has been "theres more common than not" in llm applications

Re: HyperAgents: Self-referential self-improving agents

#27
post #8

No matter how far we go, we end up with generation / discrimination architecture. Its is the core of any and all learning/exellency; exposure to chaotic perturbations allow selection of solutions that are then generalized to further, ever more straining problems; producing increasingly applicable solutions. This is the core of evolution, and is actually derivable from just a single rule.

It's a feedback loop. I've always felt that the most important part of engineering was feedback loops. Maybe nature is the greatest engineer ever?

The most important part of engineering is problem-solving, which feedback loops don't necessarily do. The reason we are here as engineers is: 2.5 billion years ago, the earth made cyanobacteria, which flourished, then flooded the earth with toxic oxygen, killing almost all life on the planet. The initial feedback loop didn't solve a problem, it destroyed a use case. That's not a solution to a problem that an engineer would choose, even if those organisms that came after were pretty happy about it...
Post reply on HN