Earlier quoted context omitted.
> DeepSeek is still a big model that requires a lot of resources to run I can run the largest model at 4 tokens per second on a 64GB card. Smaller models are _faster_ than Phi-4. I've just switched to it for my local inference.
How are you running it, can you be more specific?
Nvidia’s $589B DeepSeek rout
971–980 of 1001 posts
Re: Nvidia’s $589B DeepSeek rout
#972Earlier quoted context omitted.
Actually inference got more efficient as well, thanks to the multi-head latent attention algorithm that compresses the key-value cache to drastically reduce memory usage. https://mlnotes.substack.com/p/the-valleys-going-crazy-how-d...
If H800 is a memory-constrained model that NVIDIA built to avoid the Chinese export ban on H100 with equivalent fp8 performance, it makes zero sense to believe Elon Musk, Dario Armodei and Alexandr Wang's claims that DeepSeek smuggled H100s. The only reason why a team would allocate time on memory optimizations and writing NVPTX code rather than focusing on posttraining is if they severely struggled with memory durin…
Re: Nvidia’s $589B DeepSeek rout
#973Earlier quoted context omitted.
Can you guys explain what this would be bad for the OpenAI and Anthropic of the world? Wasn't the story always outlined to be we build better and better models, then we eventually get to AGI, AGI works on building better and better models even faster, and we eventually get to super AGI, which can work on building better and better models even faster... Isn't "super-optimization"(in the widest sense) what we expect to…
Because building a frontier model is expensive. But building a model as good as an existing frontier model is cheap (re: distillation). https://en.m.wikipedia.org/wiki/Knowledge_distillation So the takeaway is they have no moat
Re: Nvidia’s $589B DeepSeek rout
#974Still overvalued IMO. Their market cap remains ludicrous.
crickets
That demand is going to dry up any day now, mark my words. Tomorrow, even!
Re: Nvidia’s $589B DeepSeek rout
#975I feel there’s a gap missing in this thread (or I may be the one missing it) DeepSeek proved knowledge distillation works very well and cheaply https://en.m.wikipedia.org/wiki/Knowledge_distillation But they didn’t show how to build a new frontier model cheaply. So, you still need massive investments to build new frontier models. But the bad part, is they can be replicated cheaply
I think you are missing it: https://stratechery.com/2025/deepseek-faq/ That has a great overview - this is a new model, but also a distillation. They used new techniques to make it really cheap (comparatively).
Re: Nvidia’s $589B DeepSeek rout
#976I feel there’s a gap missing in this thread (or I may be the one missing it) DeepSeek proved knowledge distillation works very well and cheaply https://en.m.wikipedia.org/wiki/Knowledge_distillation But they didn’t show how to build a new frontier model cheaply. So, you still need massive investments to build new frontier models. But the bad part, is they can be replicated cheaply
This comment seems to be complete nonsense. See here https://arxiv.org/abs/2412.19437v1
Re: Nvidia’s $589B DeepSeek rout
#977Earlier quoted context omitted.
> AI doesn't seem to be one of those things where society as a whole will say, "we have enough of that; we don't need any more". Really? Has anyone made a useful, commercially successful product with it yet?
ChatGPT has over $10 million paying subscriber. No I am not counting the people using the API programmatically
The same is happening in enterprise tier products, Copilot 365 is still an extra SKU to count while Google Gemini Advanced has been integrated into the Workspace offering (i.e. they actually force you for an upsell of ~20% per user license for something we didn't ask, but I digress). At least that's a better alternative that paying +20 USD per license.
Prices need to and will go down, and business models will have to change and they are already doing so. But I'm not sure if OpenAI is really ready for that.
Re: Nvidia’s $589B DeepSeek rout
#978Earlier quoted context omitted.
The limit is high quality data, not compute.
Right and LLMs will not be able to generate their own high quality training data. There are no perpetual motion machines.
My understanding is that the whole point of R1 is that it was surprisingly effective to train on synthetic data AND to reinforce on the output rather than the whole chain of thought. Which does not require so much human-curated data and is a big part of where the efficiency gain came from.
Re: Nvidia’s $589B DeepSeek rout
#979Earlier quoted context omitted.
It has been clear for a while that one of two things is true. 1) AI stuff isn't really worth trillions, in which case Nvidia is overvalued. 2) AI stuff is really worth trillions, in which case there will be no moat, because you can cross any moat for that amount of money, e.g. you could recreate CUDA from scratch for far less than a trillion dollars and in fact Nvidia didn't spend anywhere near that much to create it…
AMD is failing to recreate CUDA. Intel is dying. Who will step in as credible competition, and when? A trillion dollar question.
Intel's fab is in trouble, but that's not the relevant part of Intel for this. They get a CUDA competitor going with GPUs built on TSMC and they're off to the races. Also, Intel's fab might very well get bailed out by the government and in the process leave them with more resources to dedicate to this.
Then you have Apple, Google, Amazon, Microsoft, any one of which have the resources to do this and they all have a reason to try.
Which isn't even considering what happens if they team up. Suppose AMD is useless at software but Google isn't and then Google does the software and releases it to the public because they're tired of paying Nvidia's margins. Suppose the whole rest of the industry gets behind an open standard.
A lot of things can happen and there's a lot of money to make them happen.
Re: Nvidia’s $589B DeepSeek rout
#980Here’s a take I haven’t seen yet: If training and inference just got 40x more efficient, but OpenAI and co. still have the same compute resources, once they’ve baked in all the DeepSeek improvements, we’re about to find out very quickly whether 40x the compute delivers 40x the performance / output quality, or if output quality has ceased to be compute-bound.