Anyone has a link to a report of its capabilities? I can't find a reliable source.
https://x.com/davis7/status/2091285712566140986
Wenghi is behind DeepSWE, one of the best benchmarks.
81–90 of 151 posts
Anyone has a link to a report of its capabilities? I can't find a reliable source.
https://x.com/davis7/status/2091285712566140986
Wenghi is behind DeepSWE, one of the best benchmarks.
I had Ox Alpha working on coding tasks for a couple days non-stop, via OpenRouter and OpenCode Zen. Crush harness. It was able to complete tasks at a level that I'd put between Sonnet and Opus. It makes few mistakes, but is not that smart. The main issue for me, is that it degraded into a doom loop several times. One of them was running the same bash command about a thousand times. The last model I've used that had t…
I usually see doom loops when working with quants. Likely theyre trying to maximize the viability of a efficient model quant that can bw upgraded. Like cutting coke to get crack, quantiry over quality.
I had Ox Alpha working on coding tasks for a couple days non-stop, via OpenRouter and OpenCode Zen. Crush harness. It was able to complete tasks at a level that I'd put between Sonnet and Opus. It makes few mistakes, but is not that smart. The main issue for me, is that it degraded into a doom loop several times. One of them was running the same bash command about a thousand times. The last model I've used that had t…
I usually see doom loops when working with quants. Likely theyre trying to maximize the viability of a efficient model quant that can bw upgraded. Like cutting coke to get crack, quantiry over quality.
Mixed signals, here it's performing below even GPT-5.4 Nano: https://livebench.ai/ while here it outperforms Fable by a significant margin: https://oxalpha.com/ but if the latter is true, will people still say it was "distilled" from Fable?
Inference was atrocious in terms of speed and constant timeouts. If it's served fast it will be a delight to use.
I had Ox Alpha working on coding tasks for a couple days non-stop, via OpenRouter and OpenCode Zen. Crush harness. It was able to complete tasks at a level that I'd put between Sonnet and Opus. It makes few mistakes, but is not that smart. The main issue for me, is that it degraded into a doom loop several times. One of them was running the same bash command about a thousand times. The last model I've used that had t…
Mixed signals, here it's performing below even GPT-5.4 Nano: https://livebench.ai/ while here it outperforms Fable by a significant margin: https://oxalpha.com/ but if the latter is true, will people still say it was "distilled" from Fable?
Claims about Ox Alpha performing at Fable level were from the social media hype cycle. Everything new in the LLM space brings a wave of influencers hyping it up as a revolutionary leap forward. Don’t forget to like and subscribe to learn more. It is a capable small model, but it’s not frontier level. The interesting part will be seeing the model size, how it responds to quantization, and how fast it runs on the kind…
Earlier quoted context omitted.
Two potentials from my pov: 1. Just variance in pass@K. If you prompt any model multiple times you'll see a large variance. N=1, but I find chinese open source models have a higher variance than higher-RL'd models like fable/opus. 2. They legitimately shipped a new RL checkpoint over the 7 days, which I find hard to believe. I am leaning towards 1.
3. Deployment problems unrelated to the weights causing degraded performance
Earlier quoted context omitted.
I really want to see hard evidence of distillation before I buy into it. Seems like a lot of sour grapes over not having the sort of lead assumed. In this field, it has been shown repeatedly that leaps in performance come swiftly and without notice.
FWIW, the way GLM-5.2 (and 5.3) talk is clearly claude, so it is for sure also trained using distillation. The metric used there is me screaming at my screen per operating hours. Does it matter? IMO not really. Weights are open after all. (Or.. soon at least for 5.3)
And critically, like contracts in general, Anthropic's terms of service is only binding upon the user/counterparty. So even if a company say specifically sought out 'claude-like' content, and claude code traces available on the internet, if they don't use the Anthropic platform there is no ToS claim.