Live data from Hacker News

Ox-Alpha Is GLM?

dejan.ai

31–40 of 75 posts

Re: Ox-Alpha Is GLM?

#31
post #18

It's not a good model tbh, got a bunch of things wrong that Opus corrected in my codebase.

Yet to find a model that cross-model review doesn’t find a bunch of things wrong with. I’m running simultaneous review with whichever of Grok4.6/GLM5.3/Fable/Sol didn’t write it, and each model tends to find items the others didn’t.

Wasn't just a review, it failed the task I gave and Opus completed the task

Re: Ox-Alpha Is GLM?

#32
post #25

I wonder if the NCD metric says something about distillation too. Would you expect that a model that has been distilled/seen traces from other models would have a smaller NCD? It would be really interesting to see if this holds up and provides evidence of distillation or certainly evidence of model outputs being used in the training mix.

https://eqbench.com/results/creative-writing-v3/hybrid_parsi...

Re: Ox-Alpha Is GLM?

#33

I think within 12 months we’re going to see a frontier (inc open models) that’s so good at almost all human-directed tasks that which model you use just won’t matter. Only differences that remain will be in deep research or very long-range tasks.

People were saying this last year, and they’ll be saying the exact same thing next year. The goalpost keeps moving.

Someone else having been too early on a prediction has little bearing on my prediction. A year ago almost nobody was using open models as daily drivers, today they are. When I run out of Fable and Sol credits in a week, I switch to GLM5.3, and it's not quite there, but it's good enough for productive work.

Re: Ox-Alpha Is GLM?

#34
post #17

If it’s not zhipu then why is it returning errors that zhipu does for other models? Who else would return the exact same errors even if they took a lot of core infra like tokenizer from z?

Someone could've trained model on top of GLM. Same way Cognition trained their SWE model on top of Kimi and Cursor did same with their Composer model.

Re: Ox-Alpha Is GLM?

#35
post #34
post #17

If it’s not zhipu then why is it returning errors that zhipu does for other models? Who else would return the exact same errors even if they took a lot of core infra like tokenizer from z?

Someone could've trained model on top of GLM. Same way Cognition trained their SWE model on top of Kimi and Cursor did same with their Composer model.

While possible the amount of variation in serving infrastructure is unlikely to land with actually giving the exact same errors zhipu does.

It feels like glm flash, and there was a report zhipu had secured a huge new cluster suggesting they have the capacity. My guess anyway.

https://www.tomshardware.com/tech-industry/artificial-intell...

Re: Ox-Alpha Is GLM?

#36
post #18

It's not a good model tbh, got a bunch of things wrong that Opus corrected in my codebase.

Yet to find a model that cross-model review doesn’t find a bunch of things wrong with. I’m running simultaneous review with whichever of Grok4.6/GLM5.3/Fable/Sol didn’t write it, and each model tends to find items the others didn’t.

If your changes are non trivial even the same model will loop over and over with the feedback.

Re: Ox-Alpha Is GLM?

#38
post #34
post #17

If it’s not zhipu then why is it returning errors that zhipu does for other models? Who else would return the exact same errors even if they took a lot of core infra like tokenizer from z?

Someone could've trained model on top of GLM. Same way Cognition trained their SWE model on top of Kimi and Cursor did same with their Composer model.

The reasoning levels are the same as GLM 5.3. GLM 5.3 is still not open...

I believe it's GLM 5.3 Flash or Air.

Re: Ox-Alpha Is GLM?

#40
post #18

It's not a good model tbh, got a bunch of things wrong that Opus corrected in my codebase.

The harness is making a big difference, lackluster performance with pi but somehow very good performance with opencode. There’s some rl there for sure, for a smaller model it’s likely going to perform much better in a harness it understands the best.
Post reply on HN