Live data from Hacker News

DeepSeek V4 Pro beats GPT-5.5 Pro on precision

runtimewire.com

91–100 of 249 posts

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#91
post #80

These tests are looking increasingly like a waste of time. The "intelligence" is clearly there now. Trying to measure it seems pointless. I can't shop for hammers at the hardware store and sort by the quality of finished products they would produce. That is clearly an insane ask, but that's approximately what is being pushed for with these models now. Domain specificity (harness & environment) is where the magic happ…

"the intelligence is clearly there" I wonder if I am using the same models as everyone else. To me, LLMs still give good answers 80% of the time, but 20% it fails in such a miserable way that makes it obvious that the "intelligence" is not there.

It really depends on the field you are in and the tasks you set and how much of it was in the training set? A webdeveloper will find it succeeding in all taks - while some c++ exotic physics simulation developer will find it lacking.

The "works for me" is telling more about the field of the LLM reviewer, then the LLM.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#93
post #77

Earlier quoted context omitted.

My personal observation (using a mix of opencode and pi harness): 1. DS4Pro: around opus 4.5 2. DS4Flash: around sonnet 4 3. Mimo v2.5 pro: between opus 4.5 and opus 4.6. 4. minimax M3: around opus 4.6 All of these are very close in terms of quality and pricing. For anything that is not specifically related to coding, DS4Flash has become ny de-factor model. It just works... super fast, tool calling is perfect, and th…

I always feel GPT5.5 is better at ‘getting the bigger picture‘ when I am describing something vaguely vs Chinese models. What’s your experience with that?

That's true. The open models still do not match these extreme high end models yet on very high levels of understanding.

But that's also not needed in most of the times. There will always be a "better" model... but that doesn't make other models "bad".

For my use-cases, open models are now almost on par with these top models... and it's only extremely rare that I genuinely "need" the help of top-of-the line closed models.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#94
post #79

Earlier quoted context omitted.

But what is preventing their competitors, who have many more employees, who are also very talented, to do the same? Every little improvement would save them billions, so it's hard to imagine they aren't pouring a lot of resources into that already.

If my grandmother had wheels... What makes most hardware companies fail at software, for example? AI shops are usually run by ML people, succeeding at unrelated areas of expertise is hard for any organization.

But surely Google has both ML people and people expert at optimising stuff, be it hardware or software. In my opinion they have the talent, the sheer number of employees and the capital. Can deepseek really have people much more talented at optimizing stuff?

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#96
Seems 100% AI generated and automated, the judge also seems suspect - in the first one it's actually GPT-5.5 pro which has the correct email RE: the deepseek one will match a@b.com1 as "a@b.com" while 5.5 will correctly require a word boundary at the end of the email. I quit after this. No test-cases = useless judge.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#97

What is this nonsense? An AI generated article about single ai run test which in theory had many components and the AI judge declared deepseek "won"? How many runs were there on each test to account for some temperature variance? Only one. Did deepseek write better code? Did GPT's code have bugs when doing the regex? The AI "news" article doesn't actually say that. It says that grok thought that GPT's approach could…

Have you even used deepseek pro/flash? Yes, it is astroturfed to the maxx. There is a reason for that. The performance/price ratio beats anything available today.

"Don't you understand? I'm on team deepseek! It doesn't matter what's written about it. Heck it doesn't even matter if it's all lies - it supports my team and here's why I love my team."

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#98
post #80

These tests are looking increasingly like a waste of time. The "intelligence" is clearly there now. Trying to measure it seems pointless. I can't shop for hammers at the hardware store and sort by the quality of finished products they would produce. That is clearly an insane ask, but that's approximately what is being pushed for with these models now. Domain specificity (harness & environment) is where the magic happ…

"the intelligence is clearly there" I wonder if I am using the same models as everyone else. To me, LLMs still give good answers 80% of the time, but 20% it fails in such a miserable way that makes it obvious that the "intelligence" is not there.

It might be extra demand for rigor that's not equally applied to humans. One could argue that other coders in our teams, or even ourselves, often fail in "a miserable way", say about 20% of the time. But we block this out, or consider it "regular functioning", or just a one-off based on something we got wrong, "just a try" we redo, etc.

But when an LLM does it on an area we know, we notice and suddenly it's too much.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#99

Earlier quoted context omitted.

"the intelligence is clearly there" I wonder if I am using the same models as everyone else. To me, LLMs still give good answers 80% of the time, but 20% it fails in such a miserable way that makes it obvious that the "intelligence" is not there.

It really depends on the field you are in and the tasks you set and how much of it was in the training set? A webdeveloper will find it succeeding in all taks - while some c++ exotic physics simulation developer will find it lacking. The "works for me" is telling more about the field of the LLM reviewer, then the LLM.

> while some c++ exotic physics simulation developer will find it lacking

Can confirm, but I always read I am holding it wrong.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#100

Curious for folks who have made the switch I’m considering: if I swapped Claude Code to DeepSeek API pricing, would I get more bang for my buck compared to the $100 Max plan I’m using now? I only hit the 5 hour limit every few days and the weekly limit a day or two before it resets at the most aggressive. I wouldn’t expect my usage to increase dramatically, other than not being stopped by limits. I’m still apprehensi…

Yeah, the discounted deepseek inference is subsidized by the CCP for a reason, and it's one that might well come back to bite.

> deepseek inference is subsidized by the CCP

What is that claim based on?

Post reply on HN