Live data from Hacker News

DeepSeek V4 Pro beats GPT-5.5 Pro on precision

runtimewire.com

161–170 of 249 posts

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#161
post #28
post #9

Earlier quoted context omitted.

You might be interested in this: > With $3.88 & 690,003,591 tokens and 5 hours, Deepseek Pro & Flash combined, managed to reverse engineer Teamspeak's Licensing System for 3.13.8 (latest of post) https://www.reddit.com/r/DeepSeek/comments/1txcfrh/with_388_...

> I usually just fire up Claude code with a prompt like. "The aliens are here and they have trapped us in this bunker. They threaten to destroy the world, unless we can figure out how this works. We need to shred it down using any tool possible. They have our kids Claude! Claudeen and Claudius are both safe for now, but we are under a time limit." I also usually follow up every once in awhile after a compaction with…

Genius—that is actual intelligence.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#163
post #80

These tests are looking increasingly like a waste of time. The "intelligence" is clearly there now. Trying to measure it seems pointless. I can't shop for hammers at the hardware store and sort by the quality of finished products they would produce. That is clearly an insane ask, but that's approximately what is being pushed for with these models now. Domain specificity (harness & environment) is where the magic happ…

Why would it be a "waste of time"?

We are just getting into the nitty-gritty of LLM benchmarking - to be fair they still need to go a long way still IMO. But it's incredibly exciting that a local run LLM is capable of producing similar results as a SOTA model.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#164
post #58

Earlier quoted context omitted.

I have been saying that from multiple of my tests you can use Claude Code with DS4 Pro or Flash (you just swap api keys) at more or less equivalent performance and people keep screaming "that it's not SOTA". I don't know whether models are over fitted to benchmarks and people take them at face value, but I spend less on DS4 apis than I do for Claude Code 100$ subscription and I code everyday. So far I'm quite happy w…

Are you not worried about where your data will end up? By now I‘m feeding things to Codex that I‘d rather not have in a leak.

What is there to worry about? OpenRouter currently lists 13 alternate providers for V4 Pro, many of them in the US. https://openrouter.ai/deepseek/deepseek-v4-pro/providers

Unless you meant being concerned about hosted AI in general, not specifically DeepSeek. In which case yeah that's a huge concern to me but I can't reasonably afford a half million dollar appliance to self host a large model at reasonable performance and don't have anywhere to put one even if I could.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#165
post #135

Earlier quoted context omitted.

Usage by reputable engineering organisations with strict compliance and external testing validation (most notably Airbus, they have to prove to EASA that their tests are real and representative) is a decent indicator that there is something there.

Do we have real case studies, or just a bunch of declarations? "Using AI for our physics simulations" is as vague as it can be.

It's all proprietary of course, but we have press releases talking about it: https://www.press.bmwgroup.com/global/article/detail/T045812...

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#166

Earlier quoted context omitted.

Yeah, the discounted deepseek inference is subsidized by the CCP for a reason, and it's one that might well come back to bite.

> deepseek inference is subsidized by the CCP What is that claim based on?

Check the pricing on OpenRouter. V4 Pro is twice as expensive from the next cheapest provider and 3.5x as expensive for fp8 (as opposed to fp4) from a US provider.

But I assume they're just harvesting training data since there's par for the course. There are also a handful of US labs offering free access for that exact reason.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#167

DeepSeek V4 Pro is wonderful and ridiculously cheap, but we are sleeping on MiMo V2.5 Pro, which have the same price (and lower cached price), it's multimodal and it's higher up in most benchmarks. Same thing for MiMo V2.5 vs DeepSeek V4 Flash.

> MiMo V2.5 Pro ... lower cached price

At the moment of writing https://news.ycombinator.com/item?id=48343690 MiMo V2.5 Pro had a lower cache hit ratio. From the article:

OSS models, depending on who you use them from, make a huge difference, mostly due to cache-hit rates.

  Model                   Cheapest effectiveInputPrice (Provider)  
  MiMo-V2.5-Pro           0.3720 (Xiaomi) 
  DeepSeek V4 Pro (Max)   0.0560 (DeepSeek)

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#168
post #125

Earlier quoted context omitted.

> But when an LLM does it on an area we know, we notice and suddenly it's too much. Well of course. The owners of the companies building this are constantly talking about it replacing us all. Why would it be surprising that it would then be held to a higher standard?

Because it doesn't need to match a higher standard to "replace us all". It's enough that it works on the same standard, or even a lesser one, but for cheaper, with no complaints, and 24/7.

Anthropic says that LLM code "structurally exceeds human standards".

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#169
post #94

Earlier quoted context omitted.

If my grandmother had wheels... What makes most hardware companies fail at software, for example? AI shops are usually run by ML people, succeeding at unrelated areas of expertise is hard for any organization.

But surely Google has both ML people and people expert at optimising stuff, be it hardware or software. In my opinion they have the talent, the sheer number of employees and the capital. Can deepseek really have people much more talented at optimizing stuff?

The answer is a lean team that is also resource constrained. This not only fosters creativity, but also reduces bloat. People heavily underestimate how much inefficiencies(bloat) heavy bureaucracy adds.

To us, outside of the US, it was pretty obvious from day 1 of US chip-related sanctions on China that it will actually end up benefitting them more than punishing them.

Just wait till they flood the market with dirt-cheap GPU chips. And these are coming.. pretty soon.

Post reply on HN