Live data from Hacker News

DeepSeek v4

api-docs.deepseek.com

341–350 of 1001 posts

Re: DeepSeek v4

#341

There are quite a few comments here about benchmark and coding performance. I would like to offer some opinions regarding its capacity for mathematics problems in an active research setting. I have a collection of novel probability and statistics problems at the masters and PhD level with varying degrees of feasibility. My test suite involves running these problems through first (often with about 2-6 papers for conte…

Wondering how gpt 5.5 is doing in your test. Happy to hear that DeepSeek has good performance in your test, because my experience seems to correlate with yours, for the coding problems I am working on. Claude doesn't seem to be so good if you stray away from writing http handlers (the modern web app stack in its various incarnations).

Re: DeepSeek v4

#342
post #311

Earlier quoted context omitted.

As a Brit I'm here for it to be honest, I'm tired of America with everything that's going on. China is not perfect but a bit of competition is healthy and needed

"not perfect" is a _very_ big simplification of what China is though

they compare it to fascist USA though

Re: DeepSeek v4

#343

Earlier quoted context omitted.

But remember to not ask about Taiwan!

Just ask it for a summary of the USA’s role in Iran, Gaza, Lebanon and its recent threats against Panama, Cuba and Greenland! It might be able to keep track.

Are you implying that western models were manipulated to hide and distort those events, like they do with the Tiananmen Square event, and Taiwan?

Re: DeepSeek v4

#345

Earlier quoted context omitted.

You can use deepseek with Claude code

You can , but does it work well? I assume CC has all kinds of Claude specific prompts in it, wouldn't you be better with a harness designed to be model agnostic like pi.dev or OpenCode?

I've been using all Kimi K2.6, gpt-5.4 and now Deepseek v4 (thought not extensively yet) in Claude Code and I can say it works much better than you'd expect. It looks like the system prompt and tools are pulling a lot of weight. Maybe the current models are good enough that you don't need them to be trained for a specific harness.

Re: DeepSeek v4

#346
lots of great stuff, but the plot in the paper is just chart crime. different shades of gray for references where sometimes you see 4 models and sometimes 3.

Re: DeepSeek v4

#347

Earlier quoted context omitted.

I'm pretty sure OpenAI and Anthropic are overpricing their token billed API usage mainly as an incentive to commit to get their subscriptions instead.

Anthropic recently dropped all inclusive use from new enterprise subscriptions, your seat sub gets you a seat with no usage. All usage is then charged at API rates. It’s like a worst of both worlds!

What's the point then? Special conditions for data retention/non-training policies?

Re: DeepSeek v4

#348

Earlier quoted context omitted.

I'm pretty sure OpenAI and Anthropic are overpricing their token billed API usage mainly as an incentive to commit to get their subscriptions instead.

The target audience for the APIs is third party apps which are not compatible with the subscriptions.

True. I missed that.

Re: DeepSeek v4

#350

At this point 'frontier model release' is a monthly cadence, Kimi 2.6 Claude 4.6 GPT 5.5, the interesting question is which evals will still be meaningful in 6 months.

more like weekly or almost daily, gpt 5.5 was literally 12 hours ago
Post reply on HN