Live data from Hacker News

The impact of competition and DeepSeek on Nvidia

youtubetranscriptoptimizer.com

321–330 of 500 posts

Re: The impact of competition and DeepSeek on Nvidia

#321

Even if DeepSeek has figured out how to do more (or at least as much) with less, doesn't the Jevons Paradox come into play? GPU sales would actually increase because even smaller companies would get the idea that they can compete in a space that only 6 months ago we assumed would be the realm of the large mega tech companies (the Metas, Googles, OpenAIs) since the small players couldn't afford to compete. Now that st…

My interpretation is that yes in the long haul, lower energy/hardware requirements might increase demand rather than decrease it. But right now, DeepSeek has demonstrated that the current bottleneck to progress is _not_ compute, which decreases the near term pressure on buying GPUs at any cost, which decreases NVIDIA's stock price.

Short term, I 100% agree, but remains to be seen what "short" means. According to at least some benchmarks, Deepseek is two full orders of magnitude cheaper for comparable performance. Massive. But that opens the door for much more elaborate "architectures" (chain of thought, architect/editor, multiple choice) etc, since it's possible to run it over and over to get better results, so raw speed & latency will still matter.

Re: The impact of competition and DeepSeek on Nvidia

#322

Even if DeepSeek has figured out how to do more (or at least as much) with less, doesn't the Jevons Paradox come into play? GPU sales would actually increase because even smaller companies would get the idea that they can compete in a space that only 6 months ago we assumed would be the realm of the large mega tech companies (the Metas, Googles, OpenAIs) since the small players couldn't afford to compete. Now that st…

It does, but proving that it can be done with cheaper (and more importantly for NVidia), lower margin chips breaks the spell that NVidia will just be eating everybody's lunch until the end of time.

Re: The impact of competition and DeepSeek on Nvidia

#323
post #318
post #174

This is a humble and informed acrticle (comparing to others written by financial analysts the past a few days). But still have the flaw of over-estimating efficiency of deploying a 687B MoE model on commodity hardware (to use locally, cloud providers will do efficient batching and it is different): you cannot do that on any single Apple hardware (need to at least hook up 2 M2 Ultra). You can barely deploy that on des…

Is it really 37B different parameters for each token? Even with the "multi-token prediction system" that the article mentions?

I don't think anyone uses MTP for inference right now. Even if you use MTP for drafting, you need to batching in the next round to "verify" it is the right token, if that happens you need to activate more experts.

DELETED: If you don't use MTP for drafting, and use MTP to skip generations, sure. But you also need to evaluate your use case to make sure you don't get penalized for doing that. Their evaluation in the paper don't use MTP for generation.

EDIT: Actually, you cannot use MTP other than drafting because you need to fill in these KV caches. So, during generation, you cannot save your compute with MTP (you save memory bandwidth, but this is more complicated for MoE model due to more activated experts).

Re: The impact of competition and DeepSeek on Nvidia

#324

Great article. > Now, you still want to train the best model you can by cleverly leveraging as much compute as you can and as many trillion tokens of high quality training data as possible, but that's just the beginning of the story in this new world; now, you could easily use incredibly huge amounts of compute just to do inference from these models at a very high level of confidence or when trying to solve extremely…

> NVIDIAs moat Offtopic, but your comment finally pushed me over the edge to semantic satiation [1] regarding the word "moat". It is incredible how this word turned up a short while ago and now it seems to be a key ingredient of every second comment. [1] https://en.wikipedia.org/wiki/Semantic_satiation

The word moat was first used in english in the 15th century https://www.merriam-webster.com/dictionary/moat

Re: The impact of competition and DeepSeek on Nvidia

#325
post #311

Earlier quoted context omitted.

The price data I (we?) get is 15 minute delayed. I would guess most of the profiteering is from consumers not knowing the last transaction prices? I.e. an artificially created edge by the broker who then sells the API to clean their hands of the scam.

Real-time price data is indeed not free, but widely available even in retail brokerages. I've never seen a 15 minute delay in any US based trade, and I think I can even access level 2 data a limited number of times on most exchanges (not that it does me much good as a retail investor). > I would guess most of the profiteering is from consumers not knowing the last transaction prices? No, not at all. And I wouldn't ev…

OK ty I guess I got it wrong. I thought it was way more common than for my scrappy bank.

Re: The impact of competition and DeepSeek on Nvidia

#326
post #322

Even if DeepSeek has figured out how to do more (or at least as much) with less, doesn't the Jevons Paradox come into play? GPU sales would actually increase because even smaller companies would get the idea that they can compete in a space that only 6 months ago we assumed would be the realm of the large mega tech companies (the Metas, Googles, OpenAIs) since the small players couldn't afford to compete. Now that st…

It does, but proving that it can be done with cheaper (and more importantly for NVidia), lower margin chips breaks the spell that NVidia will just be eating everybody's lunch until the end of time.

If demand for AI chips will increase due to Jevon’s paradox, why would Nvidia’s chips become cheaper?

In the long run, yes, they will be cheaper due to more competition and better tech. But next month? It will be more expensive.

Re: The impact of competition and DeepSeek on Nvidia

#327
post #264

The description of DeepSeek reminds me of my experience in networking in the late 80s - early 90s. Back then a really big motivator for Asynchronous Transfer Mode (ATM) and fiber-to-the-home was the promise of video on demand, which was a huge market in comparison to the Internet of the day. Just about all the work in this area ignored the potential of advanced video coding algorithms, and assumed that broadcast TV-q…

I worked on a network that used a protocol very similar to ATM (actually it was the first Iridium satellite network). An internet based on ATM would have been amazing. You’re basically guaranteeing a virtual switched circuit, instead of the packets we have today. The horror of packet switching is all the buffering it needs, since it doesn’t guarantee circuits. Bandwidth is one thing, but the real benefit is that ATM…

Man, I saw a presentation on Iridium when I was at Motorola in the early 90s, maybe 92? Not a marketing presentation - one where an engineer was talking, and had done their own slides.

What I recall is that it was at a time when Internet folks had made enormous advances in understanding congestion behavior in computer networks, and other folks (e.g. my division of Motorola) had put a lot of time into understanding the limited burstiness you get with silence suppression for packetized voice, and these folks knew nothing about it.

Re: The impact of competition and DeepSeek on Nvidia

#328

The description of DeepSeek reminds me of my experience in networking in the late 80s - early 90s. Back then a really big motivator for Asynchronous Transfer Mode (ATM) and fiber-to-the-home was the promise of video on demand, which was a huge market in comparison to the Internet of the day. Just about all the work in this area ignored the potential of advanced video coding algorithms, and assumed that broadcast TV-q…

I love algorithms as much the next guy, but not really. DCT was developed in 1972 and has a compression ratio of 100:1. H.264 compresses 2000:1. And standard resolution (480p) is ~1/30th the resolution of 4k. --- I.e. Standard resolution with DCT is smaller than 4k with H.264. Even high-definition (720p) with DCT is only twice the bandwidth of 4k H.264. Modern compression has allowed us to add a bunch more pixels, bu…

The web didn't go from streaming 480p straight to 4k. There were a couple of intermediate jumps in pixel count that were enabled in large part by better compression. Notably, there was a time period where it was important to ensure your computer had hardware support for H.264 decode, because it was taxing on low-power CPUs to do at 1080p and you weren't going to get streamed 1080p content in any simpler, less efficient codec.

Re: The impact of competition and DeepSeek on Nvidia

#329
post #258

Earlier quoted context omitted.

Google is silently catching up fast with Gemini. They're also pursuing next gen architectures like Titan. But most importantly, the frontier of AI capabilities is shifting towards using RL at inference (thinking) time to perform tasks. Who has more data than Google there? They have a gargantuan database of queries paired with subsequent web nav, actions, follow up queries etc. Nobody can recreate this, Bing failed to…

Can you say more about using RL at inference time, ideally with a pointer to read more about it? This doesn’t fit into my mental model, in a couple of ways. The main way is right in the name: “learning” isn’t something that happens at inference time; inference is generating results from already-trained models. Perhaps you’re conflating RL with multistage (e.g. “chain of thought”) inference? Or maybe you’re talking ab…

I wasn't clear. Model weights aren't changing at inference time. I meant at inference time the model will output a sequence of thoughts and actions to perform tasks given to it by the user. For instance, to answer a question it will search the web, navigate through some sites, scroll, summarize, etc. You can model this as a game played by emitting a sequence of actions in a browser. RL is the technique you want to train this component. To scale this up you need to have a massive amount of examples of sequences of actions taken in the browser, the outcome it led to, and a label for if that outcome was desirable or not. I am saying that by recording users googling stuff and emailing each other for decades Google has this massive dataset to train their RL powered browser using agent. Deepseek proving that simple RL ca be cheaply applied to a frontier LLM and have reasoning organically emerge makes this approach more obviously viable.

Re: The impact of competition and DeepSeek on Nvidia

#330

Even if DeepSeek has figured out how to do more (or at least as much) with less, doesn't the Jevons Paradox come into play? GPU sales would actually increase because even smaller companies would get the idea that they can compete in a space that only 6 months ago we assumed would be the realm of the large mega tech companies (the Metas, Googles, OpenAIs) since the small players couldn't afford to compete. Now that st…

Jevons paradox isn't some iron law like gravity.
Post reply on HN