Earlier quoted context omitted.
My conspiracy theory is that this is the new space race, and the CCP encourages this to show the world what Chinese engineers are capable of, and tank the Anthropic/OpenAI valuation bubble as a desirable side effect.
Half of the engineers at OAI and Anthropic are Asian, I don't think China does all these for signaling.
DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
211–220 of 342 posts
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#212[flagged]
I'm not looking forward to it, them being strapped for resources provides a huge incentive to develop and release these smaller models. Even if they'd still release their models once they are able to comfortably service all potential customers via their cloud, running them locally would be almost impossible due to their size.
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#213Earlier quoted context omitted.
Two (linked) DGX Sparks would do it I guess. Though probably slowly (I'd guess 15-20 tok/sec for decode, but higher for prefill). So ~$8-9k USD at current RAM prices, substantially less if they ever (sigh) drop. Electricity use would actually be relatively modest. But it makes little to no sense as long as API prices are what they are. Except for maybe privacy reasons.
the rational in one’s mind is similar to buying expensive supercar but no driving it daily. owning a few GPUs is a lot cheaper than supercars.
I don't use it for local inference so much. I use it to learn.
I also use it as my daily driving Aarch64 development system.
Aside it's also very cool what else can be done with unified GPU memory, once you realize you have it...
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#214Now let's see Dario's price cut.
They can't compete, they have bills. By the end of the year, if they can't react, it's game over.
Maybe US clients could be a little patriotic here, but money is money. They won't give them free money forever.
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#215Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#216It’s exciting that a model scoring this high is dirt cheap. It’s also so inefficient, when they release the full performance numbers it’s not going to be good. One example, it takes about 3.6x more tokens to finish the same work as Gemini Flash 3.6.
Using tokens to evaluate models is an outdated approach. Cost per task is what matters. Not all tokens are created equal
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#217Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#218Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#219Earlier quoted context omitted.
Because reddit unironically has better decorum around usage of their upvote/downvote system than HN does. People on HN downvote objectively correct information because they don't like it 24/7. There's a reason the creator of Zig left and gave the computer version of a middle finger on the way out to HN!
You define OPs post as "objectively correct information" even though it is an unknown future event for which they provided zero evidence?
Since when have hackernews started to become toxic like stackoverflow used to be?
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#220Earlier quoted context omitted.
Very true. The verses selected come from a tool call. The LLM queries the app for relevant verses. It picks the keywords - so perhaps there’s bias there but the verses are handed to the LLM. But there are ways to control and constrain the LLMs and what the user is presented with. These are all top of mind for me and why I felt there could be a better option than asking ChatGPT directly.
How polished is the experience, I see there is an invite link on the website.