Earlier quoted context omitted.
The problem is picking between models. I do not want to spend my time switching models and trying to decipher which model should be used for what. Maybe that's just a me problem that I need to figure out.
DeepSeek v4 is honestly good enough that I'm fine throwing it at everything in my hobby projects. I guess now I'll be switching from V4 Pro-Preview to V4 Flash. My only real complaint is that they can't do images, which limits their ability to autonomously debug some kinds of issues Of course you can get more bang for your buck by being more deliberate. But that's equally true with US frontier models. You can optimiz…
DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
231–240 of 342 posts
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#232Somewhat relatedly, how do the economics for Huggingface work? They must be hosting petabytes of models and datasets by now. I have downloaded quite a few “just in case”, only to replace them with the later iteration months later. Does the file hosting actually cost peanuts when you do it yourself and the cloud has shattered my understanding of what it actually costs to deliver so much data?
File hosting is pretty cheap. Egress traffic is about $90/TB in the cloud, but around $1/TB in the real world. Storage is in the realm of $5/TB/month after adding redundancy At the scale of Huggingface, that still amounts to a lot load of money. Significantly less than if you did the same in AWS, but still a lot That said, they do have a deal with AWS to make the data available in AWS ip space. Maybe they got some ch…
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#233create a plan with SOTA, execute with this.
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#234Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#235Earlier quoted context omitted.
Hell, the US doesn't even need to act. I claim the CCP will wise up within 2 years, possibly much much sooner, and ban their own companies from open sourcing to prevent the Americans from acquiring the capabilities. Despite all the nonsense claims of China distilling US models, the reality is that the Americans absolutely do distill these free Chinese models, and distillation when full logprobs are available (i.e. yo…
You’re describing the dump and pump strategy that china has historically used across a number of industries. Stands to reason that this is what is going on.
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#236Earlier quoted context omitted.
I’ve been using v4 flash for an app I’m building [1] and it’s amazing how cost effective and good it is coming from having always used gpt, opus and sonnet models. It’s so cost effective I can offer a generous free tier since my goal isn’t to make money with it. [1] https://trysojourn.app
I get where you're coming from, and the intent to make it easier for people to find examples and verses, but there's a fine line with LLMs giving you answers, is that it's interpreting it in some form. Doesn't that run counter to prevailing ideology, that you're meant to either struggle with the materials / seek understanding yourself, or have your religious leaders interpret/receive those insights?
I think it depends, Catholics wouldn't be able to use this because the Magisterium is the ultimate authority on interpreting Scripture, so the personal interpretation isn't really needed. This is not to say that Catholics don't read the Bible, they are encouraged to do so since it deepens their faith
On the other hand, for Protestant it varies, the High Church denominations are closer to Catholics (though none of them accept the Magisterium) in terms of scripture interpretation, but the Low Church ones (like Baptists or Non-Denominational ) are more open to personal interpenetration.
Disclaimer: I'm a Catholic, so if I made a mistake here fellow Protestants, please correct me.
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#237Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#238It’s exciting that a model scoring this high is dirt cheap. It’s also so inefficient, when they release the full performance numbers it’s not going to be good. One example, it takes about 3.6x more tokens to finish the same work as Gemini Flash 3.6.
That depends. Is it also more reliable?
If two books, one big one slim, prove the same thesis, what I would be interested in is the quality of the content, not the size. There can be a measure of efficiency in "have you really thought it through", but it is clearly complex - it requires measuring how solid the reasoning is.
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#239Earlier quoted context omitted.
They announced it already, read the tech report. "For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework , using the max reasoning effort level with temperature = 1.0, top_p = 0.95."
> They announced it already I meant something I could download and run.
Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
#240Earlier quoted context omitted.
My conspiracy theory is that this is the new space race, and the CCP encourages this to show the world what Chinese engineers are capable of, and tank the Anthropic/OpenAI valuation bubble as a desirable side effect.
Oh no, 1kkk market country with top-tier research labs developing its own technology, must be evil!