Live data from Hacker News

DeepSeek V4 Pro beats GPT-5.5 Pro on precision

runtimewire.com

151–160 of 249 posts

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#151
post #80

These tests are looking increasingly like a waste of time. The "intelligence" is clearly there now. Trying to measure it seems pointless. I can't shop for hammers at the hardware store and sort by the quality of finished products they would produce. That is clearly an insane ask, but that's approximately what is being pushed for with these models now. Domain specificity (harness & environment) is where the magic happ…

"the intelligence is clearly there" I wonder if I am using the same models as everyone else. To me, LLMs still give good answers 80% of the time, but 20% it fails in such a miserable way that makes it obvious that the "intelligence" is not there.

That's a better score than I'd give my own thinking.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#152

It’s four poorly constructed arbitrary experiments which say very little about the competency of either model. The article reads like thin, auto-generated ai clickbait for nerd sniping or shilling a model. Consider the lead: > DeepSeek V4 Pro wins this head-to-head by being more exact where it matters: following instructions, matching schemas, and solving edge cases cleanly. GPT-5.5 Pro is still strong, but it gave a…

I agree, I'd rather not see AI-generated articles about AI on HN unless they're really good.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#153
post #145

Earlier quoted context omitted.

I think you've misunderstood the purpose of a lead (sic). Per Merriam-Webster [^1], a lede is: > the introductory section of a news story that is intended to entice the reader to read the full story (Emphasis mine) You may prefer more matter-of-fact phrasing, of course, but criticising a lede for attempting to achieve its goal is unjustified. [^1]: https://www.merriam-webster.com/dictionary/lede

I think the criticism is less about whether the lede is good at achieving its goal and more about whether that goal is honorable in the first place. So dismissing it on technicalities is for sure clever but also obvious and lame. The Letter/spirit thing eventually got boring. Please find better material

I apologise if using words correctly is obvious and lame.

GP is explicitly criticising the language in the lede as being unsuitably vague, hence my reply.

As to the goal of the article, I fail to see what is dishonourable about comparing LLMs. You may consider the methodology flawed, but it's a perfectly respectable goal.

Sorry, was that another technicality? I'll try to find better material, just for you.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#154

Earlier quoted context omitted.

I think you've misunderstood the purpose of a lead (sic). Per Merriam-Webster [^1], a lede is: > the introductory section of a news story that is intended to entice the reader to read the full story (Emphasis mine) You may prefer more matter-of-fact phrasing, of course, but criticising a lede for attempting to achieve its goal is unjustified. [^1]: https://www.merriam-webster.com/dictionary/lede

A 'lede' is just an intentionally differentiated spelling of 'lead'; the origin of the word is just lead . Collins dictionary defines lede: a variant spelling of lead

TIL, thank you.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#156

This benchmark draws a very different picture having GPT5.5 on the very top with 70% and DeepSeek at 8% https://deepswe.datacurve.ai

DeepSWE has been heavily criticized though. https://github.com/datacurve-ai/deep-swe/issues/21 Putting GPT 5.5 on top is the obviously correct part, but everything else about it makes very little sense.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#157

It’s four poorly constructed arbitrary experiments which say very little about the competency of either model. The article reads like thin, auto-generated ai clickbait for nerd sniping or shilling a model. Consider the lead: > DeepSeek V4 Pro wins this head-to-head by being more exact where it matters: following instructions, matching schemas, and solving edge cases cleanly. GPT-5.5 Pro is still strong, but it gave a…

I think you've misunderstood the purpose of a lead (sic). Per Merriam-Webster [^1], a lede is: > the introductory section of a news story that is intended to entice the reader to read the full story (Emphasis mine) You may prefer more matter-of-fact phrasing, of course, but criticising a lede for attempting to achieve its goal is unjustified. [^1]: https://www.merriam-webster.com/dictionary/lede

It’s the hardest part of an article if you ask me.

Filling it with slop constructs signals the reader no effort was made writing the article. So no effort should be put into reading it.

The rest of the article is equally flimsy. Great clickbait title, perhaps that is even harder than writing a lede.

I am not a native speaker :)

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#159

Earlier quoted context omitted.

"the intelligence is clearly there" I wonder if I am using the same models as everyone else. To me, LLMs still give good answers 80% of the time, but 20% it fails in such a miserable way that makes it obvious that the "intelligence" is not there.

It really depends on the field you are in and the tasks you set and how much of it was in the training set? A webdeveloper will find it succeeding in all taks - while some c++ exotic physics simulation developer will find it lacking. The "works for me" is telling more about the field of the LLM reviewer, then the LLM.

Funny you used this example :)

I'm a month and a half deep into using it to make a traffic simulator with a bespoke physics engine that has complete drivetrain, suspension, and tire kernels. Think rally sim with an arcadey super off road presentation. It also has a full (also bespoke) webtransport stack that has held up beyond my wildest dreams. The simulation itself is capable of >500k cars. That was all complete about 2 weeks ago, the remainer of the work is integrating and optimizing the (you guessed it, also bespoke) pure synthesis sound engines for drivetrain/engine/tire/collision noise, and making pixi performant enough to actually display it all.

My biggest regret is actually accepting its choice of pixi, if I would have just trusted what I knew and done my own renderer too it'd already be finished! In the meantime I'm having fun boiling down the nonlinear continuous-ish models into fitted surrogate polynomials and regime-specific closed forms. Currently using cloud credits I was given to test the library I need to accelerate this work on CDNA3/4 cards. It's so nice to make someone else's room hot for a change

I've really enjoyed the ~3 month speedrun from "he has psychosis" to "the model did everything", yet somehow the number of people having this kind of success continues to match up with where I'd rank a given dev. There just aren't that many talented people out there and an even smaller subset of them are aiming high enough with LLMs, if at all. It's a truly awesome time to not have/need a job

E: Most of my frustration is directed at OAI, they keep fucking up the cache and usage calculations. They got a grand out of me, I'm excited to see what Deepseek does for me with the same.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#160
post #98

Earlier quoted context omitted.

It might be extra demand for rigor that's not equally applied to humans. One could argue that other coders in our teams, or even ourselves, often fail in "a miserable way", say about 20% of the time. But we block this out, or consider it "regular functioning", or just a one-off based on something we got wrong, "just a try" we redo, etc. But when an LLM does it on an area we know, we notice and suddenly it's too much.

No. It is not intelligent at all to confidently assert false things you know nothing about, and humans don’t do this outside of compulsive liars. For example… A few days ago I asked ChatGPT where a Spurgeon quote came from. Response: “That quote is widely attributed to Charles Spurgeon, but pinning down an exact sermon or written source is surprisingly difficult—and that’s a red flag. Short answer There’s no well-att…

>It is not intelligent at all to confidently assert false things you know nothing about, and humans don’t do this outside of compulsive liars.

"The Dunning-Kruger effect describes a disturbing cognitive bias that afflicts us all. People with limited expertise in an area tend to overestimate how much they know—and we all have gaps in our expertise." [1]

[1] https://www.openmindmag.org/articles/david-dunning-on-expert...

Post reply on HN