QwQ: Alibaba's O1-like reasoning LLM
281–290 of 435 posts
Re: QwQ: Alibaba's O1-like reasoning LLM
#282Re: QwQ: Alibaba's O1-like reasoning LLM
#283Is o1 even that good? It's doesn't even rank first on LMArena..
Re: QwQ: Alibaba's O1-like reasoning LLM
#284Re: QwQ: Alibaba's O1-like reasoning LLM
#285It would appear to have been a U.S.-only game until now. As Eric Schmidt said in the YouTube lecture (that keeps getting pulled down), LLM's have been a rich-companies game.
Re: QwQ: Alibaba's O1-like reasoning LLM
#286Is Alibaba's LLM the "Chinese LLM"? It would appear to have been a U.S.-only game until now. As Eric Schmidt said in the YouTube lecture (that keeps getting pulled down), LLM's have been a rich-companies game.
Alibaba has been pumping out a bunch of useful models for a long time.
Re: QwQ: Alibaba's O1-like reasoning LLM
#287Is o1 even that good? It's doesn't even rank first on LMArena..
There might be some narrow band of practical problems in between what other LLMs can do and what o1 can’t, but I don’t think that really matters for most use cases, especially given how much slower it is.
Day to day, you just don’t really want to prompt a model near the limits of its capabilities, because success quickly becomes a coin flip. So if a model needs five times as long to work, it needs to dramatically expand the range of problems that can be solved reliably.
Re: QwQ: Alibaba's O1-like reasoning LLM
#288Is o1 even that good? It's doesn't even rank first on LMArena..
My experience has been that typical LLMs will have more “preamble” to what they say, easing the reader (and priming themselves autoregressively) into answers with some relevant introduction of the subject, sometimes justifying the rationale and implications behind things. But for o1, that transient period and the underlying reasoning behind things is part of OpenAI’s special sauce, and they deliberately and aggressively take steps to hide it from users.
o1 will get correct answers to hard problems more often than other models (look at the math/coding/hard subsections on the leaderboard, where anecdotal experiences aside, it is #1), and there’s a strong correlation between correctness and a high score in those domains because getting code or math “right” matters more than the justification or explanation. But in more general domains where there isn’t necessarily an objective right or wrong, I know the vibe matters a lot more to me, and that’s something o1 struggles with.
Re: QwQ: Alibaba's O1-like reasoning LLM
#289Is Alibaba's LLM the "Chinese LLM"? It would appear to have been a U.S.-only game until now. As Eric Schmidt said in the YouTube lecture (that keeps getting pulled down), LLM's have been a rich-companies game.
qwen, deepseek, yi - there have been a number of high quality, open chinese competitors
Re: QwQ: Alibaba's O1-like reasoning LLM
#290Is o1 even that good? It's doesn't even rank first on LMArena..
don’t overindex on the lmsys arena, the median evaluator is kinda mid