Ornith-1.0: self-improving open-source models for agentic coding
11–20 of 65 posts
Re: Ornith-1.0: self-improving open-source models for agentic coding
#12Previously: https://news.ycombinator.com/item?id=48709744 https://swelljoe.com/post/will-it-mythos/ : "Poor performer here, only found the one bug that almost every model found, despite its performance on other benchmarks being excellent for its size. […] It also performs poorly in a chat without tools, exhibiting an ehthusiasm for hallucination. I’m currently working on a replication of this with full tool access, i…
How is that a serious phrase in '26? I mean I have no idea if this fine-tune is good, haven't tried it, but testing a (clearly) agentic model without tool access and expecting it to work is crazy, no? What was he even testing?!
Re: Ornith-1.0: self-improving open-source models for agentic coding
#13Re: Ornith-1.0: self-improving open-source models for agentic coding
#14Previously: https://news.ycombinator.com/item?id=48709744 https://swelljoe.com/post/will-it-mythos/ : "Poor performer here, only found the one bug that almost every model found, despite its performance on other benchmarks being excellent for its size. […] It also performs poorly in a chat without tools, exhibiting an ehthusiasm for hallucination. I’m currently working on a replication of this with full tool access, i…
> It also performs poorly in a chat without tools, exhibiting an ehthusiasm for hallucination. I’m currently working on a replication of this with full tool access, including bash/Python, which may allow this model to be competitive. How is that a serious phrase in '26? I mean I have no idea if this fine-tune is good, haven't tried it, but testing a (clearly) agentic model without tool access and expecting it to work…
Re: Ornith-1.0: self-improving open-source models for agentic coding
#15These are simply benchmaxxed versions of either Qwen or Gemma 4.
Re: Ornith-1.0: self-improving open-source models for agentic coding
#16Re: Ornith-1.0: self-improving open-source models for agentic coding
#17This is the first Qwen fine-tune that is not immediately rejected by the local LLM community, and in some cases even being recommended. Based on my limited usage, it is good, gives creative solutions to coding problems. I don't expect 9-35B models to one-click create full apps. Most people who were complaining did so .
It has been this way since the beginning, unfortunately. There is certainly no harm in trying on local models on local workloads with modest guardrails.
Like most of these models (Qwen, Gemma, Llama, gpt-oss), finding all the little gotchas like, special tokens and prompt structure, model preference are a PITA right now. The reward are really nice models that run exceptionally well in agentic harnesses tuned with the prompts and parameters you fought so hard to learn.
Re: Ornith-1.0: self-improving open-source models for agentic coding
#18This is the first Qwen fine-tune that is not immediately rejected by the local LLM community, and in some cases even being recommended. Based on my limited usage, it is good, gives creative solutions to coding problems. I don't expect 9-35B models to one-click create full apps. Most people who were complaining did so .
Re: Ornith-1.0: self-improving open-source models for agentic coding
#19Previously: https://news.ycombinator.com/item?id=48709744 https://swelljoe.com/post/will-it-mythos/ : "Poor performer here, only found the one bug that almost every model found, despite its performance on other benchmarks being excellent for its size. […] It also performs poorly in a chat without tools, exhibiting an ehthusiasm for hallucination. I’m currently working on a replication of this with full tool access, i…
> It also performs poorly in a chat without tools, exhibiting an ehthusiasm for hallucination. I’m currently working on a replication of this with full tool access, including bash/Python, which may allow this model to be competitive. How is that a serious phrase in '26? I mean I have no idea if this fine-tune is good, haven't tried it, but testing a (clearly) agentic model without tool access and expecting it to work…