There are quite a few comments here about benchmark and coding performance. I would like to offer some opinions regarding its capacity for mathematics problems in an active research setting. I have a collection of novel probability and statistics problems at the masters and PhD level with varying degrees of feasibility. My test suite involves running these problems through first (often with about 2-6 papers for conte…
DeepSeek v4
341–350 of 1001 posts
Re: DeepSeek v4
#342Earlier quoted context omitted.
As a Brit I'm here for it to be honest, I'm tired of America with everything that's going on. China is not perfect but a bit of competition is healthy and needed
"not perfect" is a _very_ big simplification of what China is though
Re: DeepSeek v4
#343Earlier quoted context omitted.
But remember to not ask about Taiwan!
Just ask it for a summary of the USA’s role in Iran, Gaza, Lebanon and its recent threats against Panama, Cuba and Greenland! It might be able to keep track.
Re: DeepSeek v4
#344Any way to connect this to claude code?
Re: DeepSeek v4
#345Earlier quoted context omitted.
You can use deepseek with Claude code
You can , but does it work well? I assume CC has all kinds of Claude specific prompts in it, wouldn't you be better with a harness designed to be model agnostic like pi.dev or OpenCode?
Re: DeepSeek v4
#346Re: DeepSeek v4
#347Earlier quoted context omitted.
I'm pretty sure OpenAI and Anthropic are overpricing their token billed API usage mainly as an incentive to commit to get their subscriptions instead.
Anthropic recently dropped all inclusive use from new enterprise subscriptions, your seat sub gets you a seat with no usage. All usage is then charged at API rates. It’s like a worst of both worlds!
Re: DeepSeek v4
#348Earlier quoted context omitted.
I'm pretty sure OpenAI and Anthropic are overpricing their token billed API usage mainly as an incentive to commit to get their subscriptions instead.
The target audience for the APIs is third party apps which are not compatible with the subscriptions.
Re: DeepSeek v4
#349Any way to connect this to claude code?
Re: DeepSeek v4
#350At this point 'frontier model release' is a monthly cadence, Kimi 2.6 Claude 4.6 GPT 5.5, the interesting question is which evals will still be meaningful in 6 months.