Live data from Hacker News

DeepSeek V4 Pro beats GPT-5.5 Pro on precision

runtimewire.com

141–150 of 249 posts

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#142
post #98

Earlier quoted context omitted.

"the intelligence is clearly there" I wonder if I am using the same models as everyone else. To me, LLMs still give good answers 80% of the time, but 20% it fails in such a miserable way that makes it obvious that the "intelligence" is not there.

It might be extra demand for rigor that's not equally applied to humans. One could argue that other coders in our teams, or even ourselves, often fail in "a miserable way", say about 20% of the time. But we block this out, or consider it "regular functioning", or just a one-off based on something we got wrong, "just a try" we redo, etc. But when an LLM does it on an area we know, we notice and suddenly it's too much.

No. It is not intelligent at all to confidently assert false things you know nothing about, and humans don’t do this outside of compulsive liars. For example…

A few days ago I asked ChatGPT where a Spurgeon quote came from. Response:

“That quote is widely attributed to Charles Spurgeon, but pinning down an exact sermon or written source is surprisingly difficult—and that’s a red flag.

Short answer There’s no well-attested primary source (sermon, lecture, or publication) where Spurgeon clearly says that exact wording.” Etc. etc. … Why it sounds like Spurgeon It fits his theology and rhetoric almost perfectly: • etc etc. … Closest authentic themes (but not the quote) Spurgeon repeatedly says things like: • etc etc. … So the quote is basically: a modern condensation of real Spurgeon ideas, not a verifiable citation etc. etc.”

Utter bullshit. One web search produces the full sermon manuscript with the quote.

One could argue that the previous context in the thread primed the LLM to fail here, but once again, a person is not confused by the change of topic.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#143
post #80

These tests are looking increasingly like a waste of time. The "intelligence" is clearly there now. Trying to measure it seems pointless. I can't shop for hammers at the hardware store and sort by the quality of finished products they would produce. That is clearly an insane ask, but that's approximately what is being pushed for with these models now. Domain specificity (harness & environment) is where the magic happ…

[deleted]

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#144
post #80

These tests are looking increasingly like a waste of time. The "intelligence" is clearly there now. Trying to measure it seems pointless. I can't shop for hammers at the hardware store and sort by the quality of finished products they would produce. That is clearly an insane ask, but that's approximately what is being pushed for with these models now. Domain specificity (harness & environment) is where the magic happ…

"the intelligence is clearly there" I wonder if I am using the same models as everyone else. To me, LLMs still give good answers 80% of the time, but 20% it fails in such a miserable way that makes it obvious that the "intelligence" is not there.

In my experience of hiring and managing people, I would have been very happy if they gave good answers or produced good results 80% of the time.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#145

It’s four poorly constructed arbitrary experiments which say very little about the competency of either model. The article reads like thin, auto-generated ai clickbait for nerd sniping or shilling a model. Consider the lead: > DeepSeek V4 Pro wins this head-to-head by being more exact where it matters: following instructions, matching schemas, and solving edge cases cleanly. GPT-5.5 Pro is still strong, but it gave a…

I think you've misunderstood the purpose of a lead (sic). Per Merriam-Webster [^1], a lede is: > the introductory section of a news story that is intended to entice the reader to read the full story (Emphasis mine) You may prefer more matter-of-fact phrasing, of course, but criticising a lede for attempting to achieve its goal is unjustified. [^1]: https://www.merriam-webster.com/dictionary/lede

I think the criticism is less about whether the lede is good at achieving its goal and more about whether that goal is honorable in the first place.

So dismissing it on technicalities is for sure clever but also obvious and lame.

The Letter/spirit thing eventually got boring. Please find better material

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#146
Why was this posted to HN? What an utter waste of time. Someone's slopwriter writes a slop article about which slopper slops the most slopulicious slop. Comments agree it's a bogus "study". We need some gate on AI-written articles. It's so weird that AI-written comments are not permitted, while the front page can be occupied by stuff like this.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#147

It’s four poorly constructed arbitrary experiments which say very little about the competency of either model. The article reads like thin, auto-generated ai clickbait for nerd sniping or shilling a model. Consider the lead: > DeepSeek V4 Pro wins this head-to-head by being more exact where it matters: following instructions, matching schemas, and solving edge cases cleanly. GPT-5.5 Pro is still strong, but it gave a…

I think you've misunderstood the purpose of a lead (sic). Per Merriam-Webster [^1], a lede is: > the introductory section of a news story that is intended to entice the reader to read the full story (Emphasis mine) You may prefer more matter-of-fact phrasing, of course, but criticising a lede for attempting to achieve its goal is unjustified. [^1]: https://www.merriam-webster.com/dictionary/lede

A 'lede' is just an intentionally differentiated spelling of 'lead'; the origin of the word is just lead. Collins dictionary defines lede: a variant spelling of lead

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#149

Earlier quoted context omitted.

I have both subscriptions and I definitely feel gpt is better and more consistent, but when I run out of limits I don't miss it too much

That's the whole point. The tool you have vs. the expensive tools you don't have because they're too expensive. I don't feel like paying 100 times the price for a 1-5% better tool.

The cutting edge of LLM-based software engineering seems to be all about how to harness the "good enough" pseudo-intelligence of consumer-level affordable models into achieving practical results, through iterations, tests, harnesses, etc. And these models are getting smarter every month, including open-weight models people can run on their own machines and servers. We're not seeing the kind of leaps as often as before, but it hasn't plateau'ed yet, the models are getting better all the time.

It implies that eventually open-weight models like DeepSeek, which are self-hostable locally or on premises, will become good enough for more people and businesses, in terms of productivity gains versus cost. Consumer hardware will adapt to that demand, making it even more affordable and within reach.

Not sure how that speculation fits with the billions of dollars of investment that AI companies will need to convert to profit somehow.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#150
post #4

... according to grok-4-1-fast-non-reasoning who was the judge, on 4 tasks in total, score was 38 to 33 so obviously huge conclusions can be made. > We ran 4 fresh text tasks, generated on the fly for this matchup so neither model could prepare in advance, and had grok-4-1-fast-non-reasoning score each one. DeepSeek: DeepSeek V4 Pro scored 38.0 to OpenAI: GPT-5.5 Pro's 33.0.

grok-4-1-fast was retired about a month ago. Requests to grok-4-1-fast-non-reasoning now silently route to grok-4.3 (a 5x more expensive model), with reasoning set to "none". https://docs.x.ai/developers/migration/may-15-retirement TFA was published today, which implies grok-4.3 was used.

What specific single model being used is like the least of the issues with their methodology.
Post reply on HN