Live data from Hacker News

DeepSeek V4 Pro beats GPT-5.5 Pro on precision

runtimewire.com

241–249 of 249 posts

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#241

Earlier quoted context omitted.

The only thing I could read from your posts is that you are team openai and completely mad that people are abandoning chatgpt

"You're on the team against me so I oppose everything you say". Again it's the same problem - what you're doing. I'm not on "team OpenAI". I'm also not on "team deepseek". I'm commenting on how so much of the population is literally unable to see the world unless it is filtered through some "team" lens that they are for or against. Judge the material based on what's in the material. Not as it boosting or hurting your…

> The material in this article is crap judge it as crap and say so regardless of your team.

Your area again making the same mistake as before.

You are making the most passionate defense of team openai pretending that other people are making irrational claims.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#242
post #190

Earlier quoted context omitted.

i have been using deepseek-v4-flash since it came out. i use a highly structured harness and spec/test driven workflow running through opencode, and so far there has been nothing it can't do. i have run through a bunch of tests: re-writing vvenc with assembly kernels, creating the first generation agent harness integration with opencode, porting TS npm modules to C++, porting an entire TS server app to C++, creating…

Thanks for the details. What's a second generation agent? You mentioned the workflow is heavy on specs and tests. The smaller models seem to be really good at following instructions now. (Well, some of them!) So that's probably part of why you're seeing good results. It has a very clear target. Whereas with more open ended instructions they seem to struggle more. I think common sense is the main thing you get with mo…

> What's a second generation agent?

i meant that i initially developed an agent harness as a set of skills integrated with opencode and now i am in the process of using that to write a new agent from scratch to replace opencode.

> probably part of why you're seeing good results

yes. i think tests and setting up feedback loops for diagnosing errors (logs, debugging, etc) are the most important things. in my experience deepseek-v4-flash tends to ignore instructions to use these tools and default to churning through code and guessing the cause of errors, which is often wrong, so it requires occasionally stepping in when it has been grinding fruitlessly for a while and reminding it, probably due to context length and sparse attention forgetting instructions that are put in context at the beginning of a session.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#243

Earlier quoted context omitted.

"You're on the team against me so I oppose everything you say". Again it's the same problem - what you're doing. I'm not on "team OpenAI". I'm also not on "team deepseek". I'm commenting on how so much of the population is literally unable to see the world unless it is filtered through some "team" lens that they are for or against. Judge the material based on what's in the material. Not as it boosting or hurting your…

> The material in this article is crap judge it as crap and say so regardless of your team. Your area again making the same mistake as before. You are making the most passionate defense of team openai pretending that other people are making irrational claims.

> Your area again making the same mistake as before.

> You are making the most passionate defense of team openai

At no point did I mention Openai, referr to openai or imply anything about openai (just mentioned your reference). Nothing I'm saying weighs in on any form of discussion or debate between Deepseek & Open Models vs OpenAI.

The fact that you are unable to separate those two is your failing, not mine. Your argument is the equivalent of the following:

A: Deepseek ran into a burning building last week and saved 10,000 orphans from a fire.

Me: No Deepseek did not save 10,000 orphans from a burning building last week. Regardless of what you think of Deepseek it didn't save 10,000 orphans. It's an LLM in a computer, not a humanoid robot - if you look at that for 2 seconds you see that claim is nonsense.

You: By attacking those supporting Deekseek you have declared yourself for team OpenAI and are clearly an OpenAI supporter!

Me: Saying deepseek didn't save 10k orphans has nothing to do with OpenAI. It is a lie saying that deepseek saved 10k lives. It's an LLM chat bot. Regardless of how anyone feels about deepseek - discuss it on it's merits not on bs.

You: See! You keep defending OpenAI you open AI shill! Stop passionately defending OpenAI!

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#244

Earlier quoted context omitted.

> The material in this article is crap judge it as crap and say so regardless of your team. Your area again making the same mistake as before. You are making the most passionate defense of team openai pretending that other people are making irrational claims.

> Your area again making the same mistake as before. > You are making the most passionate defense of team openai At no point did I mention Openai, referr to openai or imply anything about openai (just mentioned your reference). Nothing I'm saying weighs in on any form of discussion or debate between Deepseek & Open Models vs OpenAI. The fact that you are unable to separate those two is your failing, not mine. Your ar…

You are hallucinating, my friend.

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#246

Earlier quoted context omitted.

> MiMo V2.5 Pro ... lower cached price At the moment of writing https://news.ycombinator.com/item?id=48343690 MiMo V2.5 Pro had a lower cache hit ratio. From the article: OSS models, depending on who you use them from, make a huge difference, mostly due to cache-hit rates. Model Cheapest effectiveInputPrice (Provider) MiMo-V2.5-Pro 0.3720 (Xiaomi) DeepSeek V4 Pro (Max) 0.0560 (DeepSeek)

Could it be that it changed recently, or am I missing something? Both prices are the same https://openrouter.ai/compare/xiaomi/mimo-v2.5-pro/deepseek/... EDIT: okay I misread it, does this mean that DeepSeek reuses a higher percentage of tokens at cache price that MiMo, am I right?

Correct. According to https://minimaxir.com/2026/05/openrouter-hy3/#llm-economics-... [0] when served by DeepSeek, Cache Read Costs/Input Costs are a very low percentage:

  DeepSeek V4 Pro    0.83%
  DeepSeek V4 Flash  2%
Notice that OpenRouter response caching is not available when account-level ZDR is enforced [1]

[0] https://news.ycombinator.com/item?id=48317294#48317823 [1] https://openrouter.ai/docs/guides/features/response-caching#...

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#247
post #217

Earlier quoted context omitted.

The creation--which isn't "his" in the first place, by any standard definition--was not only itself "derived from" our creations but was always supposed to be "open".

> which isn't "his" in the first place, by any standard definition I was saying that because of the previous comment: > to Scam Altman's creation It wasn't derived in the same way though - I can read loads of books and so can write my own book, but that's not derivation in the same way as the Deepseek's derivation.

[flagged]

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#248
post #217

Earlier quoted context omitted.

The creation--which isn't "his" in the first place, by any standard definition--was not only itself "derived from" our creations but was always supposed to be "open".

> which isn't "his" in the first place, by any standard definition I was saying that because of the previous comment: > to Scam Altman's creation It wasn't derived in the same way though - I can read loads of books and so can write my own book, but that's not derivation in the same way as the Deepseek's derivation.

[flagged]

Re: DeepSeek V4 Pro beats GPT-5.5 Pro on precision

#249

Earlier quoted context omitted.

Damned squishy humans, with their feelings and moods...

Indeed. It's like saying "the strongest human on their best day can support the roof of this tent for hours, how dare you criticise them for being squishy humans" when someone says "why don't we make an a-frame out of wood?"

LLMs don't make a good A-frame, nor would I classify them as wood-like. People propose LLMs as solutions as if they're wooden when they're teetering contraptions of metal rods, aluminum extrusions, rubber bands, and duct tape. That can do the trick. It can't be relied on to fail reliably like a single solid material like wood.
Post reply on HN