Live data from Hacker News

Simple Meta-Harness on Islo.dev

zozo123.github.io

11–20 of 24 posts

Re: Simple Meta-Harness on Islo.dev

#11
post #8

serious question: I've already got a opencode harness running on a local model. It's easily installable via the insecure bash command. It's already tailored with a couple of plugins and with a proper TODO.md and planning, I can get it to loop fine with proper attention to its pratfalls on vague/non-determinant language. It's all running on a AMD 395+ Qwen3-Coder-Next model with ~256k context. opencode has a webui I c…

You know how OpenCode can be prompted to modify itself when you want to improve it in some way? This just automates that kind of thing.

It can't actually; I had to create a systemd service that watched the config path and send a signal to reload the files. It roughly works, but it doesn't actually do the loop correctly.

However, the problem with self-modification is the tendency towards inoperable states. Does it automatically revert when a detrimental state is reached? How does it determine that a modification worked?

Re: Simple Meta-Harness on Islo.dev

#12
This is not how I've seen the term meta-harness be used. The common usage I've seen has been for a meta-harness to be a wrapper around an existing agent to give that agent a new ui or abilities.

Re: Simple Meta-Harness on Islo.dev

#13
post #10

It has now become fashionable to dress oneself in the garb of science to sell dev environments ... for agents. It has now become fashionable to claim much, and furnish little. It has now become fashionable to fail to understand or state the core of your proposal in as few words as possible: instead of "genetic algorithm applied to the space of harnesses, parallelized by our infrastructure" we get "Three swaps. Same o…

we need better RL

Re: Simple Meta-Harness on Islo.dev

#14
post #8

Earlier quoted context omitted.

You know how OpenCode can be prompted to modify itself when you want to improve it in some way? This just automates that kind of thing.

It can't actually; I had to create a systemd service that watched the config path and send a signal to reload the files. It roughly works, but it doesn't actually do the loop correctly. However, the problem with self-modification is the tendency towards inoperable states. Does it automatically revert when a detrimental state is reached? How does it determine that a modification worked?

The paper shows that it can. Take note this seems to be someone’s experiment. If it’s not working for you that’s probably because it’s not a polished product.

Re: Simple Meta-Harness on Islo.dev

#15
I did this too, ablating all the components in my coding agent harness. The insight from my meta-optimization loops was "have judge agents review the plan and implementation".

One of my own insights here is that you need to collect not just execution traces, but all the human-in-the-loop nudges and steering commands. They are one of the purest sources of feedback on coding agents when seen in context.

I agree with OP on the need to collect traces and compare them, not just scores. It is a much richer source of feedback.

If anyone is interested I have a slide deck about my approach: https://horiacristescu.github.io/claude-playbook-plugin/docs...

Re: Simple Meta-Harness on Islo.dev

#17
post #4

This seems to be another over optimization for AI that many are trying to get into. The LLM's improve, and your setup is deprecated, you wasted time optimizing for a slight edge. TDLR: You trade time for slight edge.

i don't disagree, though harness engineering is a real discipline that even the best AI labs put their brightest minds on, and the loop itself doesn't deprecate when models improve.

Re: Simple Meta-Harness on Islo.dev

#19

I have no idea what this does or is. I really wish they could have given a better description of why this is useful.

Yeah I have been reading a lot of posts like this lately. Technical blog post clearly written by an LLM summarizing something vibe-coded. They always start using project-specific jargon right away and they never give you enough context or backstory to understand why this thing exists. It's seems very clearly to be a symptom of someone pointing an LLM at a repo and telling it "write a github page for this project".

It really shines through in pieces like this that LLM's have a severely constrained worldview and underdeveloped theory of mind. They can't imagine that a line like "A 200-line POC that goes from 0/5 to 5/5 in four proposer steps" means nothing to me as a subtitle for the page. After all "proposer steps" and "5/5" are *right there* in it's context. Surely everyone has "proposer steps" in their context, right?

Re: Simple Meta-Harness on Islo.dev

#20
I distilled the Meta-Harness workflow in a skill [0]. Unlike the original Islo POC, which demonstrates an automated runtime loop converging from trace-rich evaluations [1], my test only evaluates whether the distilled skill improves a lead agent's prompt-repair discipline and audit trail [2].

It took a few tries to figure out what to test in the first place, since it is not obvious what the workflow should improve (prompt? guided agent ability?).

So, the only meaningful test I ended up with was giving easy tasks, but with a deliberately misleading/incomplete prompt, then testing whether persisting deltas and observations between successive prompts meaningfully improves a meta-agent's ability to correct the imprecise prompt (what I mean by "prompt-repair discipline and audit trail") [2].

From a couple more experiments (summarized in [2]), I found that the Meta-agent does not really have an effect on how well the guided agents perform, but simply improves imprecise prompts better.

My conclusion is that this method works to improve bad prompts, I didn't demonstrate that improves guided agent capabilities. However, I think it's better to work on your prompts before giving them to agents instead of giving bad prompts and iterating on them with a meta-agent.

[0]: https://github.com/ouatu-ro/skill-distillery/blob/main/skill...

[1]: https://github.com/zozo123/meta-harness-on-islo

[2]: https://github.com/ouatu-ro/skill-distillery/blob/main/repor...

Post reply on HN