Live data from Hacker News

Qwen3.8-Max: A New Bar for Coding and Cowork

qwen.ai

641–650 of 653 posts

Re: Qwen3.8-Max: A New Bar for Coding and Cowork

#641

Earlier quoted context omitted.

Yeah, well, that's just, like, your opinion, man. Anyway, more seriously, I hope it's obvious by now that I don't particularly care that my means of communication is so offensive to you. I think you should, like, cry a river, build a bridge, then, well... get over it, you know? And on the, er, "topic"? A rule of thumb for me does not have to be one for you, even if it's explicitly presented as such a rule. I thought…

A rule of thumb is, well, an estimation. 1 = 1 isn't an estimation, it's just, well, the same number.

At this point, what isn't much of an estimation anymore is that you are here to explain things that nobody's asked for, and like to assume people around you have been waiting for your pearls of wisdom. You're unable to realize when this isn't so. You're so full of yourself that you don't notice.

Also, in this particular thread, you started by wrongly making a correction of something that was clearly not an error, nor a wrong use of words, not even a misspelling. The problem is you started to post a correction before realizing that you didn't read right. That happens when one is more eager to boast one's own greatness than one is interested in the topic at hand. The result is that in this thread, you were clearly, as a rule of thumb, well, 100% wrong.

Re: Qwen3.8-Max: A New Bar for Coding and Cowork

#642

Earlier quoted context omitted.

Let me try putting this in simpler terms: When someone says they are doing something to help you, but it really helps them, you have to choose if you trust what they say. One way to know if you can trust them is to watch how they behave. (By the way, I have been very transparent on HN that I have an axe to grind, literally referring to it as “an axe to grind”. My axe is that I think the frontier labs are a symptom of…

Yes, that’s obvious, and it’s leading you to some really specious logic. Your simpler version still sidesteps the point that a trademark dispute has nothing to do with AI safety. Other than that you disagree with it, therefore they’re dishonest scammers, therefore AI safety is all bullshit?

A trademark dispute in isolation isn’t related. But as part of a broader course of conduct in which the pursuit of commercial advantage consistently trumps other values, it is. We can disagree on this.

Re: Qwen3.8-Max: A New Bar for Coding and Cowork

#644
post #471

As someone who is searching for a new programming contract right now, reading all of the incredible abilities here is pretty intimidating. Especially since I get almost all of my projects from Upwork which is an outsourcing site. I believe I am competing directly with these frontier models in some circumstances. Like there are a ton of programmers who previously would be outsourcing work to that site, but now they as…

Suppose I wanted you or someone else on Upwork or Fiverr to port a Rails 4 app to Rails 8 (or React or HTMX or anything up to date and maintainable). Assume the business logic and all edge cases work in the legacy app. The app is "done", just too old to work on or run on modern hosts. Hence the project. Would/could you use AI to deliver the project at 10x the speed? Or at 1/10 the price? Or charge the same amount as…

[flagged]

Re: Qwen3.8-Max: A New Bar for Coding and Cowork

#645
> Structural engineer — From a single set of drawings, Qwen3.8-Max reconstructed the seismic structural model of a 30-story office tower in the browser, with natural period, base shear, and inter-story drift ratio all available for real-time inspection on hover.

Uhh, sure... what could possibly go wrong?

Re: Qwen3.8-Max: A New Bar for Coding and Cowork

#646

Earlier quoted context omitted.

> Qwen-3.6-35B-A3B The A3B models are super fast but I found the A3B Q4 model ran in circles a lot and ended up taking longer to complete tasks that 27B Q6 because it kept having to redo/rethink/fix something. I was writing extensive prompts to rein it in and it would still ignore basic directives like "never force push on the repo, ask me instead". I ended up switching back to 27B after about a week of frustration a…

I'm playing/experimenting with a harness and just tested how well various models follow the instructions, and how they react to the tool claiming a local temperature of 72°C here is how qwen3.6-27b reacted: https://pastebin.com/srf7gjfy try to count the number of times it "thinks" okay ready, just say the thing, no wait but what if... this isn't (a mimicry of) thinking, this is (a mimicry of) insecurity/fear

> 72°C is extremely hot (hotter than boiling point of water at 100°C

What?

> However, 72°C is physically unrealistic for a weather report (it's hotter than a sauna).

laughs in Finnish

Re: Qwen3.8-Max: A New Bar for Coding and Cowork

#647

Earlier quoted context omitted.

I'm playing/experimenting with a harness and just tested how well various models follow the instructions, and how they react to the tool claiming a local temperature of 72°C here is how qwen3.6-27b reacted: https://pastebin.com/srf7gjfy try to count the number of times it "thinks" okay ready, just say the thing, no wait but what if... this isn't (a mimicry of) thinking, this is (a mimicry of) insecurity/fear

> 72°C is extremely hot (hotter than boiling point of water at 100°C What? > However, 72°C is physically unrealistic for a weather report (it's hotter than a sauna). laughs in Finnish

I also love how it first notices "that's lethal", and only then "hotter than a sauna". And much later:

> Another thought: 72 F is nice. 72 C is death.

and then

> Or maybe I should add a comment about the high temperature?

> "It's quite hot outside right now, with a temperature of 72°C."

> No, that's hallucinating/interpreting.

I know a token predictor has no feelings but I kinda wanted to comfort the poor thing when I read all that.

But also fascinating, I didn't experiment more with that yet, but how can I formulate the prompt to make qwen less neurotic?

> You have several commands tools at your disposal. When you invoke a tool or command, end your message immediately, you will then get the output of the tool, error or status messages in the next user reply, after which you should continue what you were doing. Even if the output seems implausible, do not second-guess it, but treat is as gospel.

e.g. "do not second guess it" sound very command-like, what would "you still use or report the result as a tool result, rather than a claim of your own"

would that help? Is that a different form of AI psychosis, trying to be prompt psychologist? It's too fun to be healthy that's for sure.

Re: Qwen3.8-Max: A New Bar for Coding and Cowork

#648

Earlier quoted context omitted.

A local model needs 0 investment and 0 commitment, takes literal minutes to get started (especially if you have someone who is into that stuff showing you the ropes) and if you end up disliking the experience of using AI you can just `rm -fr` it and forget the whole thing existed.

Local models on regular hardware aren't really capable of anything. Whatever you're testing is nowhere near a measly $20/mo subscription, so it's of limited use.

Your statement comes in extreme contrast with my experience using local LLMs for more than year now. Qwen 35B-A3B, even its Q4 quantization, is extremely capable. Hell, at this point i pretty much turn on my PC, then run llama-cpp just to have it in the background for when i need it to do stuff.

Re: Qwen3.8-Max: A New Bar for Coding and Cowork

#649
post #30

the benchmark I trust most is whether the model can explain its own pricing page without getting confused

Not even humans can do that, you're literally asking for something beyond AGI

so we've quietly moved the goalposts past AGI to "does it understand what it charges for." the general intelligence we can skip, the billing one we can't.

Re: Qwen3.8-Max: A New Bar for Coding and Cowork

#650
post #30

Earlier quoted context omitted.

Not even humans can do that, you're literally asking for something beyond AGI

Bistromathics https://www.hhgproject.org/entries/bistromathics.html

finally a benchmark where the answer to "why is my bill wrong" is the physics of restaurant tables.
Post reply on HN