Live data from Hacker News

DeepSeek V4 Pro beats GPT-5.5 Pro on precision

runtimewire.com

191–200 of 249 posts

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#191

Earlier quoted context omitted.

> poorly constructed arbitrary experiments which say very little about the competency of either model. No one ever says this about the “pelican on a bicycle” metric

Simon's pelican is in fact routinely criticised for exactly that.

Here it is on the latest Opus release 11 days ago, it’s the 5th highest voted comment on the post and the most critical comment is “should you at least try like 10 times or something to average the random effects”:

https://news.ycombinator.com/item?id=48311979

Gemini Flash release 19 days ago, again no criticism:

https://news.ycombinator.com/item?id=48198232

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#192

I was using Claude until they banned Opencode, and now use GPT at my day job. I've been using Deepseek through Opencode Go on the $10/mo plan, and I honestly can't really tell much difference. Its just as capable, and makes the same kinds of dumb mistakes and the other two have been making since March. For the price, I'm more than happy with it.

It's interesting. 95% of time you don't need the extra 5% rigor that frontier models provide to you compared to the 10-100x cheaper Chinese equivalents. The remaining 5% of time you get a big boost for your high-reasoning problem solving needs and evade a lot of pain. Now, I just need to be able to predict accurately when I need this extra 5% and when not :)

I find the trick I use is to get the model to come up with a phased plan, and review it. If I spot anything that seems dumb, I give direction on the way it should be done. And once you finalize that, the model can run through the steps fairly reliably. As long as you're intentionally making all the big decisions, things tend to work out well.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#193
post #98

Earlier quoted context omitted.

It might be extra demand for rigor that's not equally applied to humans. One could argue that other coders in our teams, or even ourselves, often fail in "a miserable way", say about 20% of the time. But we block this out, or consider it "regular functioning", or just a one-off based on something we got wrong, "just a try" we redo, etc. But when an LLM does it on an area we know, we notice and suddenly it's too much.

Because a human fails in a known way. If a human does not have expertise in domain X or tech Y, they will fail there and the expectation is that they will fail. With an LLM you never know where it can fail. There is no domain expertise for an LLM. It can fail in a miserable way in the same domain it worked spectacularly for.

Humans fail in infinitely more complicated ways than LLMs. They can have a difficult personality, a medical issue, family stress, hangover, sleep deprivation or they can just wake on the wrong side of the bed. On any given day, you never know if you will get an expert in domain X or a sleep-deprived version of the same that accidentally drops a database.

Indeed, if you remember before AI took the world by storm, HN used to be chock-full of articles about how the hiring process is broken for both employers and candidates, where you can never tell if what you see is what you get.

When I run a local LLM I get none of that. I hit the intelligence walls or buggy behaviour, but it doesn't matter if it's 8am or 8pm, the model behaves exactly the same. If something doesn't work as I wished, I can retry as many times as I wanted without the model getting angry at me.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#194

Earlier quoted context omitted.

> poorly constructed arbitrary experiments which say very little about the competency of either model. No one ever says this about the “pelican on a bicycle” metric

I am willing to guess it is but gets downvoted or similar. Simon is a bit of a cult of personality on HN for better or worse.

I have his blog in my RSS app and I click every pelican test because it's fun. I think criticizing it for lack of scientific or technical rigor kind of misses its point. It's a fun curiosity.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#195
post #193

Earlier quoted context omitted.

Because a human fails in a known way. If a human does not have expertise in domain X or tech Y, they will fail there and the expectation is that they will fail. With an LLM you never know where it can fail. There is no domain expertise for an LLM. It can fail in a miserable way in the same domain it worked spectacularly for.

Humans fail in infinitely more complicated ways than LLMs. They can have a difficult personality, a medical issue, family stress, hangover, sleep deprivation or they can just wake on the wrong side of the bed. On any given day, you never know if you will get an expert in domain X or a sleep-deprived version of the same that accidentally drops a database. Indeed, if you remember before AI took the world by storm, HN u…

Damned squishy humans, with their feelings and moods...

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#196
post #94

Earlier quoted context omitted.

If my grandmother had wheels... What makes most hardware companies fail at software, for example? AI shops are usually run by ML people, succeeding at unrelated areas of expertise is hard for any organization.

But surely Google has both ML people and people expert at optimising stuff, be it hardware or software. In my opinion they have the talent, the sheer number of employees and the capital. Can deepseek really have people much more talented at optimizing stuff?

No I don't think they can, but then Google literally has their own custom inference hardware that they target so ... yeah 3.5 flash is extremely pricey compared to v4 pro and now I'm wondering why that would be. It's difficult to imagine they don't care given we know they're prepared to pay $2B / mo for additional GPU capacity.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#198

It’s four poorly constructed arbitrary experiments which say very little about the competency of either model. The article reads like thin, auto-generated ai clickbait for nerd sniping or shilling a model. Consider the lead: > DeepSeek V4 Pro wins this head-to-head by being more exact where it matters: following instructions, matching schemas, and solving edge cases cleanly. GPT-5.5 Pro is still strong, but it gave a…

In the car business there is only one or two car models that are the best ideal choice, but many subpar companies and models, are still selling for many reasons. It shows DeepSeek is competitive, if not better sometimes, than GPT 5.5. Also shows there is no moat. As such it is a highly significant signal.

I agree that there may be a lot of variation between models that leads to different use cases, at least today. But I’m not sure the car analogy works.

An X5 is not simply “inferior” to a CR-V, or vice versa. A Camry is not “inferior” to an F-150, or vice versa. They are optimized for different buyers, budgets, constraints, and use cases.

That may actually be the better analogy for AI models: there probably is not one universal “best” model. There are models that are better or worse for particular tasks, price points, latency requirements, deployment constraints, privacy needs, etc.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#199

Earlier quoted context omitted.

No. It is not intelligent at all to confidently assert false things you know nothing about, and humans don’t do this outside of compulsive liars. For example… A few days ago I asked ChatGPT where a Spurgeon quote came from. Response: “That quote is widely attributed to Charles Spurgeon, but pinning down an exact sermon or written source is surprisingly difficult—and that’s a red flag. Short answer There’s no well-att…

>It is not intelligent at all to confidently assert false things you know nothing about, and humans don’t do this outside of compulsive liars. "The Dunning-Kruger effect describes a disturbing cognitive bias that afflicts us all. People with limited expertise in an area tend to overestimate how much they know—and we all have gaps in our expertise." [1] [1] https://www.openmindmag.org/articles/david-dunning-on-expert.…

Doubting if a random quote is correct is understandable given how often the training data has explanations that random quotes from famous people aren’t real. But it isn’t intelligent to proclaim that when you have the internet as a resource.

Nobody that I know would do this.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#200
post #80

These tests are looking increasingly like a waste of time. The "intelligence" is clearly there now. Trying to measure it seems pointless. I can't shop for hammers at the hardware store and sort by the quality of finished products they would produce. That is clearly an insane ask, but that's approximately what is being pushed for with these models now. Domain specificity (harness & environment) is where the magic happ…

"the intelligence is clearly there" I wonder if I am using the same models as everyone else. To me, LLMs still give good answers 80% of the time, but 20% it fails in such a miserable way that makes it obvious that the "intelligence" is not there.

[dead]
Post reply on HN