Ornith-1.0: Self-scaffolding LLMs for agentic coding
deep-reinforce.com
Ornith-1.0: Self-scaffolding LLMs for agentic coding
1–10 of 11 posts
Re: Ornith-1.0: Self-scaffolding LLMs for agentic coding
#2Re: Ornith-1.0: Self-scaffolding LLMs for agentic coding
#3I added this to a benchmark I've been doing of how well agents find security bugs, specifically security bugs originally found by Mythos. It performs poorly with only read/grep/ls tools, but in a follow-up test with a full shell and Python, it doubled its findings (still a poor showing, but it does at least indicate it is doing what it says on the tin: making tools to help it solve problems). It also did worse than Q…
Re: Ornith-1.0: Self-scaffolding LLMs for agentic coding
#4Re: Ornith-1.0: Self-scaffolding LLMs for agentic coding
#5If that is the case, this isn't just a fancy way to perform prompt optimization?
Re: Ornith-1.0: Self-scaffolding LLMs for agentic coding
#6I'd have expected this to get more HN attention. Qwen 3.6 35B capability in a 9B model is a bonkers claim.
Re: Ornith-1.0: Self-scaffolding LLMs for agentic coding
#7I'd have expected this to get more HN attention. Qwen 3.6 35B capability in a 9B model is a bonkers claim.
In my brief tests, Ornith 35B performed quite well. It won't replace DeepSeek V4 Flash for me, but if it was fast and cheap enough it might.
I don't remember being super impressed with Ornith 9B, but I could see it being on par with Qwen 3.5 35B.
Re: Ornith-1.0: Self-scaffolding LLMs for agentic coding
#8I added this to a benchmark I've been doing of how well agents find security bugs, specifically security bugs originally found by Mythos. It performs poorly with only read/grep/ls tools, but in a follow-up test with a full shell and Python, it doubled its findings (still a poor showing, but it does at least indicate it is doing what it says on the tin: making tools to help it solve problems). It also did worse than Q…
Re: Ornith-1.0: Self-scaffolding LLMs for agentic coding
#9I'd have expected this to get more HN attention. Qwen 3.6 35B capability in a 9B model is a bonkers claim.
It looks like they're comparing Orinth 9B to Qwen 3.5 35B, not Qwen 3.6. I guess it kind of makes sense since it's a finetune of 3.5, but I totally missed until I looked closely. In my brief tests, Ornith 35B performed quite well. It won't replace DeepSeek V4 Flash for me, but if it was fast and cheap enough it might. I don't remember being super impressed with Ornith 9B, but I could see it being on par with Qwen 3.5…
Re: Ornith-1.0: Self-scaffolding LLMs for agentic coding
#10I'd have expected this to get more HN attention. Qwen 3.6 35B capability in a 9B model is a bonkers claim.
It was fairly good at diagnosing the bugs once informed of their symptoms. However, if I mischaracterized the symptoms, it would weigh my input too heavily and reject its own (correct) hunch about the root cause.
So it's an interesting one. There's definitely some latent capability in there that arguably exceeds Qwen 3.6, which is absolutely no small feat. But that capability seems to come in a somewhat erratic package.
It's probably worth benchmarking it unquantized if you can. I've grown to suspect that quantization damages small models more than perplexity and KL divergence accurately reflect.
I'll also give the 9B weights a shot when I can.