Live data from Hacker News

Qwen3: Think deeper, act faster

qwenlm.github.io

291–300 of 412 posts

Re: Qwen3: Think deeper, act faster

#291
post #227

Earlier quoted context omitted.

> That would literally change the art world forever. In what world? Some small percentage up or who knows, and _that_ revolutionized art? Not a few years ago, but now, this. Wow.

Forever, as in for a few weeks… ;-)

oh boy I had a smirk after reading this comment because its partially true.

When deepseek r1 came, it lit the markets on fire (atleast american) and then many thought it would be the best forever / for a long time.

Then came grok3 , then claude 3.7 , then gemini 2.5 pro.

Now people comment that gemini 2.5 pro is going to stay forever. When deepseek came, there were articles like this on HN: "Of course, open source is the future of AI" When Gemini 2.5 Pro came there were articles like this: "Of course, google build its own gpu's , and they had the deepnet which specialized in reinforced learning, Of course they were going to go to the Top"

We as humans are just trying to justify why certain company built something more powerful than other companies. But the fact is, that AI is still a black box, People were literally say for llama 4:

"I think llama 4 is going to be the best open source model, Zuck doesn't like to lose"

Nothing is forever, its all opinions and current benchmarks. We want the best thing in benchmark and then we want an even better thing, and we would justify why / how that better thing was built.

Every time, I saw a new model rise, people used to say it would be forever.

And every time, Something new beat to it and people forgot the last time somebody said something like forever.

So yea, deepseek r1 -> grok 3 -> claude 3.7 -> gemini 2.5 pro (Current state of the art?), each transition was just some weeks IIRC.

Your comment is a literal fact that people of AI forget.

Re: Qwen3: Think deeper, act faster

#292
post #274
post #256

Earlier quoted context omitted.

Personally (anecdata) I haven't experienced any practical progress in my day-to-day tasks for a long time, no matter how good they became at gaming the benchmarks. They keep being impressive at what they're good at (aggregating sources to solve a very well known problem) and terrible at what they're bad at (actually thinking through novel problems or old problems with few sources). E.g. all ChatGPT, Claude and Gemini…

Absolutely. All models ar terrible with Objective-C and Swift, compared to let's say JS/HTML/Python. However, I've realized that Claude Code is extremely useful for generating somewhat simple landing pages for some of my projects. It spits out static html+js which is easy to host, with somewhat good looking design. The code isn't the best and to some extent isn't maintainable by a human at all, but it gets the job do…

Interesting, I'll have to try that. All the "static" page generators I've tried require React....

Re: Qwen3: Think deeper, act faster

#293
post #274

Earlier quoted context omitted.

Absolutely. All models ar terrible with Objective-C and Swift, compared to let's say JS/HTML/Python. However, I've realized that Claude Code is extremely useful for generating somewhat simple landing pages for some of my projects. It spits out static html+js which is easy to host, with somewhat good looking design. The code isn't the best and to some extent isn't maintainable by a human at all, but it gets the job do…

Building a basic static html landing page is ridiculously easy though. What js is even needed? If it's just an html file and maybe a stylesheet of course it's easy to host. You can apply 20 lines of css and have a decent looking page. These aren't hard problems.

> These aren't hard problems.

So why do so many LLMs fail at them?

Re: Qwen3: Think deeper, act faster

#294

Earlier quoted context omitted.

I don't really want it added to the training set, but eh. Here you go: > Assume I have a 3D printer that's currently printing, and I pause the print. What expends more energy, keeping the hotend at some temperature above room temperature and heating it up the rest of the way when I want to use it, or turning it completely off and then heat it all the way when I need it? Is there an amount of time beyond which the ans…

What kind of answer do you expect? It all depends on the hotend shape and material, temperature differences, how fast air moves in the room, humidity of the air, etc.

Keeping something above room temperature will always use more energy than letting it cool down and heating it back up when needed

Re: Qwen3: Think deeper, act faster

#295
post #270
post #256

Earlier quoted context omitted.

Personally (anecdata) I haven't experienced any practical progress in my day-to-day tasks for a long time, no matter how good they became at gaming the benchmarks. They keep being impressive at what they're good at (aggregating sources to solve a very well known problem) and terrible at what they're bad at (actually thinking through novel problems or old problems with few sources). E.g. all ChatGPT, Claude and Gemini…

Absolutely, as soon as they hit that mark where things get really specialized, they start failing a lot. They do generalizations on well documented areas pretty good. I only use it for getting a second opinion as it can search through a lot of documents quickly and find me alternative means.

They have broad knowledge, a lot of it, and they work fast. That should be a useful combination-

And indeed it is. Essentially every time I buy something these days, I use Deep Research (Gemini 2.5) to first make a shortlist of options. It’s great at that, and often it also points out issues I wouldn’t have thought about.

Leave the final decisions to a super slow / smart intelligence (a human), by all means, but for people who claim that LLMs are useless I can only conclude that they haven’t tried very hard.

Re: Qwen3: Think deeper, act faster

#296
post #93

Earlier quoted context omitted.

Alibaba, I have a huge favor to ask if you're listening. You guys very obviously care about the community. We need an answer to gpt-image-1. Can you please pair Qwen with Wan? That would literally change the art world forever. gpt-image-1 is an almost wholesale replacement of ComfyUI and SD/Flux ControlNets. I can't underscore how big of a deal it is. As such, OpenAI has leapt ahead and threatens to start capturing m…

I don't know, the AI image quality has gotten good but it's still slop. We are forgetting what makes art, well art. I am not even an artist but yeah I see people using AI for photos and they were so horrendous pre chatgpt-imagen that I had literally told one person if you are going to use AI images, might as well use chatgpt for it. Also though I would also like to get something like chatgpt-image generating qualitie…

Even Katy Perry started using AI for her tour backdrop visuals and it looks... well, horrendous https://twitter.com/bklynb4by/status/1915514396421337171

Re: Qwen3: Think deeper, act faster

#297
post #25

Any news on some viable successor of LLMs that could take us to AGI? As I see they still can't solve some fundamental stuff to make it really work in any scenario (halucinations, reasoning, grounding in reality, updating long-term memory, etc.)

AGIs probably comes from neurosymbolic AI. But LLMs could be the neuro-part of that. On the other hand, LLM progress feels like bullshit, gaming benchmarks and other problems occured. So either in two years all hail our AGI/AMI (machine intelligence) overlords, or the bubble bursts.

Amusingly enough, people writing stuff like the above, to my mind come over as doing what they are accusing LLMs of doing. :-)

And in discussions "is it or isn't it, AI smarter than HI already", reminds me to "remember how 'smart' an average HI is, then remember half are to the left of that center". :-O

Re: Qwen3: Think deeper, act faster

#298

With all the different open-weight models appearing, is there some way of figuring out what model would work with sensible speed (> X tok/s) on a standard desktop GPU ? I.e. I have Quadro RTX 4000 with 8G vram and seeing all the models https://ollama.com/search here with all the different sizes, I am absolutely at loss which models with which sizes would be fast enough. I.e. there is no point of me downloading the la…

Mozilla started LocalScore for exactly what you're looking for: https://www.localscore.ai/

Fascinating that 5090 is often close but not quite as good as 4090 and RTX 6000 ADA. Perhaps it indicates that 5090 has those infamous missing computational units?

3090Ti seems to hold up quite well.

Re: Qwen3: Think deeper, act faster

#299

Earlier quoted context omitted.

I similarly have a small, simple spatial reasoning problem that only reasoning models get right, and not all of them, and which Qwen3 on max reasoning still gets wrong. > I put a coin in a cup and slam it upside-down on a glass table. I can't see the coin because the cup is over it. I slide a mirror under the table and see heads. What will I see if I take the cup (and the mirror) away?

Sonnet 3.7 non-reasoning got it right. I'll think this through step by step. When you place a coin in a cup and slam it upside-down on a glass table, the coin will be between the table surface and the cup. When you look at the reflection in the mirror beneath the table, you're seeing the bottom side of the coin through the glass. Since the mirror shows heads, you're seeing the heads side of the coin reflected in the…

Not reasoning mode, but I struggle to call that “non-reasoning”.

Re: Qwen3: Think deeper, act faster

#300

Earlier quoted context omitted.

Except, very literally, data is a collection of single points (ie what we call "anecdotes").

No, Wittgenstein's rule following paradox, Shannon sampling theorem, the law that infinite polynomials pass through any finite set of points (does that have a name?), etc, etc. are all equivalent at the limit to the idea that no amount of anecdotes-per-se add up to anything other than coincidence

Anti-realism, indeterminancy, intuitionism, and radical subjectivity are extremely unpopular opinions here. Folks here are to dense to imagine that the cogito is fake bullshit and wrong. You're fighting an extremely uphill battle.

Paul Feyerabend is spinning in his grave.

Post reply on HN